
Writer has released Palmyra X6, a new AI model built on an open source foundation, combined with infrastructure upgrades it says will cut customer costs by as much as 50% for basic tasks.
The move reflects broader enterprise pressure to control AI spending, and Writer's research suggests that optimizing the infrastructure layer (the harness) can be more effective at reducing costs than switching models alone.
What happened
Writer, an AI platform for marketers, released Palmyra X6—a new flagship model built on Z.ai's open source GLM-5.2—alongside major upgrades to its infrastructure harness. The company estimates the combination will cut deployment costs for customers by as much as 50% for basic tasks, available to clients starting Thursday.
Why it matters
Enterprise customers are increasingly focused on controlling AI costs rather than chasing performance benchmarks. Writer's internal research found that harness optimization (the infrastructure layer) reduced costs by an average of 40% across multiple models—often more effective than model choice alone—suggesting that infrastructure efficiency, not just model selection, is key to cost control.
What to watch
Writer's approach keeps Palmyra X6 model-agnostic, allowing it to sit alongside other Writer models or external models imported through Azure or Amazon Bedrock, giving customers flexibility to mix and match. CEO May Habib also flagged growing enterprise skepticism toward major AI labs, which have financial incentives to increase token usage.
Writer, a provider of AI tools and agents for marketers, announced on Thursday the launch of Palmyra X6, a new flagship model designed to address enterprise cost concerns in AI deployment. The model is built as a post-training variation of Z.ai's open source GLM-5.2, and Writer estimates that Palmyra X6 combined with infrastructure harness upgrades will cut costs for customers by as much as 50% for basic tasks. Both the new model and the harness enhancements became available to Writer clients immediately.
The company's approach reflects a broader industry shift: as customers become increasingly conscious of deployment costs, the focus has moved away from simply adopting the most capable models toward optimizing for cost and efficiency. To support this shift, Writer also released significant upgrades to its standard agentic harness—the infrastructure layer that executes AI tasks. The emphasis is on enabling complex, multi-step tasks to run faster and consume fewer tokens.
Writer's internal research bolsters this strategy. A recent paper from Writer researchers tested small changes in harness efficiency across multiple different models and found that, in many cases, changes to the harness were a more reliable way to reduce costs than model choice itself, with costs falling an average of 40% across their testing. As the researchers wrote, "The harness is the one component whose efficiency multiplies across every model an organization runs—present and future." This finding suggests that infrastructure optimization may be undervalued relative to model selection in enterprise cost reduction.
Writer's approach maintains flexibility for its clients: Palmyra X6 will sit alongside other Writer models or external models imported through Azure or Amazon Bedrock. CEO May Habib framed the shift as a response to enterprise frustration. "I think the enterprise is absolutely sick of chasing the next benchmark," he told TechCrunch. "They want flattening cost, and it seems like nobody can deliver that." Habib also noted a broader erosion of confidence in major AI labs among enterprise CIOs, attributing this partly to the labs' financial incentive to drive up token use, suggesting that enterprises see them as misaligned with the goal of cost control.
Enterprise customers are facing what CEO May Habib calls an "unprecedented" cost explosion from AI deployments, prompting a shift away from the benchmark-chasing mentality that has dominated the industry. Rather than simply adopting larger or more capable models, organizations now prioritize controlling per-token costs and total deployment spend. Writer's response—combining a purpose-built model with infrastructure optimization—reflects a broader realization that model choice alone is not the main lever. Internal research by Writer's team demonstrated that harness efficiency (the optimization of how a system executes tasks) can yield more consistent and substantial cost reductions than model selection, with an average 40% reduction across multiple models tested. This suggests that the infrastructure layer, which multiplies its efficiency gains across every model an organization runs, may be more critical than previously acknowledged. Habib also notes a growing skepticism among enterprise CIOs toward major AI labs, attributing this partly to the labs' financial incentive to increase token consumption—a dynamic that may accelerate demand for alternative approaches like Writer's.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
DeepSeek, a Chinese AI startup, has officially launched V4 Pro with enhanced agent capabilities while signific…

Alibaba Cloud has begun commercial service for its Zhenwu M890 supernode in Ulanqab, Inner Mongolia, offering…

SMIC is considering breaking out revenue from AI-related peripheral chips as a standalone line item in its fin…

Airbnb shares rose on Thursday following a strong second-quarter report showing revenue up 17% year over year…

Michael Burry criticized Nvidia's memorandum of understanding with six major asset managers—Apollo Global Mana…

Nvidia CEO Jensen Huang is partnering with major Wall Street firms — Goldman Sachs, BlackRock, Blackstone, KKR…

The AI news that matters, in one minute each morning.
Sign up free