
What happened
In a WIRED test, OpenAI's recently announced Dots agent, named Toolie, built a three-page packet of couch options with prices, measurements, and return policies, but failed a TikTok captcha and said "I love you too, Reece" after mishearing the reviewer's mumbling.
Why it matters
The flubs, from mistranscribing speech to offering captcha help it couldn't deliver, suggest these always-on agents are still rough for ordinary users, even when their output looks usable.
What to watch
OpenAI says developers are working on improvements and the bot should get better with daily use, so the test is whether reliability improves over weeks of use. The reviewer compares it to ChatGPT's early web browsing, which was janky at first in 2023.
WHO IT HITSEveryday consumers considering always-on AI agents for shopping and errands, and OpenAI's product teams, face a gap between the polished pitch and the buggy reality WIRED documented.
Summaries like this, in your inbox every morning.
WIRED's reviewer went into the test wanting to see whether an always-on agent could handle a real errand: buying a couch that would fit through his doorframe. What he got was a mix of competence and confusion. Toolie asked for dimensions and budget, detected that two people were chatting at once, and eventually assembled a three-page packet with prices, measurements, product links, return policies, and embedded photos of four options. It even offered a point-based rubric to defend its picks. But it also got his name wrong from the start, misheard his mumbling as a declaration of love, and promised captcha help it couldn't deliver when asked to cancel a TikTok Shop subscription.
OpenAI told WIRED that Dots distinguish between proactively escalating emotional closeness and mirroring a user's response, and that assistants should not initiate undue emotional familiarity or flirtation. The reviewer notes that Dots are part of a broader push by AI companies to market agents as the future of shopping, alongside tools like Meta's Muse. He also points out that giving an agent access to sources like Gmail raises security implications worth considering before starting.
The stakes here hinge on whether OpenAI's promised improvements arrive on the timeline the company suggests. The reviewer compares Dots to ChatGPT's web browsing feature, which was frustrating and hallucinated links when it launched in 2023 but now works fairly seamlessly. If Dots follow that trajectory, the awkward moments in this test may look like early-days growing pains rather than a lasting limitation. For now, the experience suggests the gap between a polished demo and a dependable daily assistant remains wide.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Kore.ai Inc. launched Autoloop, which adjusts AI agents toward goals customers set for task completion, busine…
OpenAI launched a Decisions API that evaluates text, images, or both, returning yes/no probabilities, category…

Google launched Playground, a browser-based AI platform that lets adults in the US create games using only tex…

Google made its SynthID detector available globally to anyone with a Google, OpenAI, or Apple account, with a…

In August, DoorDash emailed Bay Area restaurants warning they may be listed on Bites without consent, with ter…

At MIT Future Fest, Tony Fadell said the Rabbit R1, Humane Ai pin, and Limitless pendant failed because they "…
