
What happened
Cognition is applying GPT-6 Astra across Devin, its CLI and desktop products. In one example, Devin tested an iPhone game, Otter Run, and returned a simulator recording plus a report.
Why it matters
As Cognition's teams write more code, reviewing it became a challenge. Yan says Astra's ability to test and prove its work is one of the big improvements over prior models.
What to watch
Yan says Cognition expects to manually look at less code over time and ship more, so the test is whether Devin's recordings and reports reduce manual code examination. Customer bug screenshots are already fixed and returned as screenshots.
WHO IT HITSEngineering teams that use Devin for software work stand to spend less time reading code line by line, though the body frames that as Cognition's expectation rather than a measured result.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Cognition's Devin is an autonomous software engineer used by businesses ranging from big banks to tech-native startups. As Cognition's own engineering teams write more code, reviewing that output has become a challenge, and the company is pointing to GPT-6 Astra's testing ability as the way to make that review more efficient.
The move is not limited to Devin. Cognition says it is applying GPT-6 Astra across its product lineup, including the core cloud agent Devin, its CLI and its desktop products. One illustration: Devin used Astra to test Otter Run, an iPhone game, returning a recording of the game running in a simulator alongside a report identifying checks that passed and areas left untested. The recording shows how the application behaves, while the report documents the scope of the testing, so engineers can inspect the software and see what still needs attention. Astra is also helping Cognition get back to customers faster, according to co-founder Walden Yan: when a customer sends a screenshot of a bug, the team can pass it to Devin using Astra, which fixes the issue and returns a screenshot showing the result.
What the outcome hinges on is whether Devin's testing evidence actually cuts the manual examination engineers do. Yan frames this as an expectation rather than a measured result, saying the company expects to manually look at less code over time and ship more. For engineering teams using Devin, the practical test is likely whether those recordings and reports are enough to trust a change without reading the code line by line.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…

Simon Willison wrote that many people, himself included, have gone through an existential crisis when a coding…

Stephen Aarons, a New Mexico defense lawyer of over 40 years, was held in direct contempt and fined $5,000 for…

Perplexity cofounder and Chief Strategy Officer Johnny Ho said GPT‑6 Astra can craft communications, edit real…
