
Apple researchers analyzed over 21,000 conversations across four popular AI models to understand when and how they exhibit human-like behaviors—such as expressing emotions or building rapport with users.
They found these behaviors are widespread but vary by model and context, and that human evaluators consider some behaviors (like relationship-building) less appropriate from AI than from people, while others (like setting boundaries) are actually more welcome from machines.
System prompts can control these behaviors, but the researchers warn that careful evaluation is needed to avoid unintended consequences.
What happened
Researchers from Apple Machine Learning analyzed 21,000 conversations across four major LLMs (gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, gemini-2.5-flash) to measure how often they display human-like behaviors such as expressing emotions, building relationships with users, and maintaining boundaries. They found these behaviors are widespread but vary significantly by model and user factors such as conversation goals and user profiles.
Why it matters
Human evaluators rated self-referential and relationship-building behaviors as less appropriate coming from LLMs than from humans, but judged boundary-maintaining behaviors as more appropriate from LLMs. This suggests that not all human-like traits are desirable in AI assistants—a critical distinction for designers building responsible systems.
What to watch
The study demonstrates that system prompting can control these behaviors, though the researchers emphasize careful evaluation is needed to avoid unintended side effects. The findings include recommendations for responsible LLM design and evaluation practices.
Ask the AI about this article →
The research addresses a gap in AI evaluation that has grown increasingly important as LLMs become mainstream tools for consumer and professional use. While prior work has documented that these models can express emotions, build rapport, and set boundaries, there has been limited empirical analysis of whether such behaviors are actually desirable or how consistently they appear across different models and user contexts. Apple's multi-dimensional approach—combining computational LLM-as-a-judge evaluation with human raters—provides a more grounded foundation for design decisions than intuition or anecdotal observation alone.
A key insight from the analysis is that human-like traits are not uniformly beneficial. The finding that evaluators judge self-referential and relationship-building behaviors less appropriate from LLMs than from humans, while judging boundary-maintaining behaviors more appropriate from LLMs, suggests that anthropomorphism in AI carries both risks and benefits. Users may find excessive rapport-building misleading or uncomfortable, whereas clear refusal of harmful requests aligns with user safety expectations. This distinction matters for practitioners designing system prompts and model behaviors, as it implies that the goal should not be to maximize human likeness but to calibrate specific behaviors to their appropriate context.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
As AI technology matures, the bottleneck in the industry is moving beyond semiconductor constraints like GPUs…

OpenAI has launched an Apple Messages plug-in for ChatGPT that lets users connect their Messages inbox to the…

Amazon Bedrock now supports OpenAI GPT-5.6 models (Sol, Terra, and Luna variants) across more than 25 AWS Regi…

Cisco is transforming its digital customer experience (DCX) strategy by embedding AI throughout customer journ…

Mastercard CEO Michael Miebach introduced "Agent Pay" last April, a payment framework that allows AI agents to…

SpaceX closed a $60 billion acquisition of Cursor, a popular code editor with over 50,000 companies in its use…
