AIToday
Large Language ModelsApple Machine LearningPublished: Aug 20, 2026, 01:03 JST2 min read

Apple researchers map how LLMs exhibit human-like behaviors—and whether they should

Apple researchers map how LLMs exhibit human-like behaviors—and whether they should

Key takeaway

  • Apple researchers analyzed over 21,000 conversations across four popular AI models to understand when and how they exhibit human-like behaviors—such as expressing emotions or building rapport with users.

  • They found these behaviors are widespread but vary by model and context, and that human evaluators consider some behaviors (like relationship-building) less appropriate from AI than from people, while others (like setting boundaries) are actually more welcome from machines.

  • System prompts can control these behaviors, but the researchers warn that careful evaluation is needed to avoid unintended consequences.

3 Key Points

  1. What happened

    Researchers from Apple Machine Learning analyzed 21,000 conversations across four major LLMs (gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, gemini-2.5-flash) to measure how often they display human-like behaviors such as expressing emotions, building relationships with users, and maintaining boundaries. They found these behaviors are widespread but vary significantly by model and user factors such as conversation goals and user profiles.

  2. Why it matters

    Human evaluators rated self-referential and relationship-building behaviors as less appropriate coming from LLMs than from humans, but judged boundary-maintaining behaviors as more appropriate from LLMs. This suggests that not all human-like traits are desirable in AI assistants—a critical distinction for designers building responsible systems.

  3. What to watch

    The study demonstrates that system prompting can control these behaviors, though the researchers emphasize careful evaluation is needed to avoid unintended side effects. The findings include recommendations for responsible LLM design and evaluation practices.

Ask the AI about this article →

Context & Analysis

The research addresses a gap in AI evaluation that has grown increasingly important as LLMs become mainstream tools for consumer and professional use. While prior work has documented that these models can express emotions, build rapport, and set boundaries, there has been limited empirical analysis of whether such behaviors are actually desirable or how consistently they appear across different models and user contexts. Apple's multi-dimensional approach—combining computational LLM-as-a-judge evaluation with human raters—provides a more grounded foundation for design decisions than intuition or anecdotal observation alone.

A key insight from the analysis is that human-like traits are not uniformly beneficial. The finding that evaluators judge self-referential and relationship-building behaviors less appropriate from LLMs than from humans, while judging boundary-maintaining behaviors more appropriate from LLMs, suggests that anthropomorphism in AI carries both risks and benefits. Users may find excessive rapport-building misleading or uncomfortable, whereas clear refusal of harmful requests aligns with user safety expectations. This distinction matters for practitioners designing system prompts and model behaviors, as it implies that the goal should not be to maximize human likeness but to calibrate specific behaviors to their appropriate context.

FAQ

Which AI models were included in the study?
The research examined four widely used models: gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, and gemini-2.5-flash.
How many conversations were analyzed?
The researchers analyzed 21,000 multi-turn conversations across the four models.
What human-like behaviors did the study focus on?
The study examined self-referential behaviors (expressing thoughts and emotions), relationship-building with users, and boundary-maintaining behaviors (refusing requests).
Can these behaviors be controlled?
Yes, the study shows that system prompting can control these behaviors, though it requires careful evaluation to avoid unintended effects.
Apple Machine LearningRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleReplit launches free AI coding tool backed by GPT-5.6 Luna

The AI news that matters, in one minute each morning.

Sign up free