AITodayYour daily AI briefing

Large Language Models

Jul 19, 2026

Large Language Models

The Gist

Alphabet's delayed release of Gemini 3.5 Pro has rattled investors, while competitors like Anthropic's Claude are gaining ground by focusing on practical tools rather than just raw model power. OpenAI's ChatGPT advertising platform is expanding into Japan, signaling how AI is reshaping marketing, while open-source models are proving competitive on specialized tasks like medical exams. The AI industry continues to shift focus from building bigger models to creating better applications and understanding what makes these systems actually useful.

Today's Stories

  1. 1

    Gemini 3.5 Pro delayed months; Alphabet shares slip 15% from peak

    Alphabet said in May that Gemini 3.5 Pro would launch in June, but the flagship AI model remains in testing months behind schedule. Bloomberg reported that Google updated training data to improve coding skills, with disappointing results. The company has not announced a new launch date. Coding is the main battleground for AI labs and a core service enterprise customers pay for. Alphabet is spending as much as $190 billion(約30兆円) on AI infrastructure this year, betting on Gemini staying competitive at the frontier. A delayed model is concerning given rivals OpenAI and Anthropic are shipping models that some inside Google worry have passed Gemini by.

    Alphabet reports second-quarter results on Wednesday, July 22. Key metrics include Google Cloud's growth rate and any launch timing management reveals for Gemini 3.5 Pro. So far, the business appears healthy—Google Cloud revenue jumped 63% year over year to $20.0 billion(約3.2兆円) in Q1, and the company's cloud backlog nearly doubled from the prior quarter to over $460 billion(約74兆円).

  2. 2

    Claude Code's Real Edge: The Harness, Not the Model

    A developer rebuilt Claude Code's agent architecture in CrewAI, an open-source framework, and found that the gap between a basic agent loop and Claude Code's capability comes from the surrounding machinery—planning, memory, sandboxing, subagent delegation, and approval systems—not the underlying language model itself. Most teams underestimate how much engineering sits outside the model. A bare agent loop fails on real codebases (reads wrong files, loses context, fills memory with stale output), while Claude Code stays on track. Understanding this split—model as "brain" deciding actions, harness as "hands" executing them—shows what you actually need to build reliable coding agents in production.

    The rebuild tested against a small BankAccount class with two real bugs and five tests. The harness took the project from 3 failing and 2 passing to all 5 passing, demonstrating that planning, subagents, and sandboxing together enable the agent to fix code correctly without shortcuts like editing tests.

  3. 3

    Open-weight AI models pass Swedish medical exam via training method

    Researchers demonstrated that open-weight large language models (AI systems that understand and generate text) can pass the Swedish medical licensing exam using two training techniques — supervised fine-tuning (SFT) and reinforcement learning from verification rewards (RLVR). Open-weight models (whose internal workings are publicly available, unlike proprietary systems) have historically lagged behind closed commercial models on specialized professional tasks. This result shows that with the right training approach, publicly available models can reach professional-level performance on high-stakes medical knowledge, which could make advanced AI capabilities more accessible beyond the companies that control proprietary systems.

    This is a research finding posted to an academic community (Reddit's r/MachineLearning). The body does not specify which model was used, the exact pass rate, how this compares to prior benchmarks, or any timeline for practical deployment in medical education or licensing.

  4. 4

    Reddit CS student debates: deep coding skills vs. AI-era pivot

    A Computer Science student in Pakistan solicited advice on Reddit about whether to pursue traditional software engineering skills (Java, backend, system design, algorithms) or pivot toward AI and automation, citing their brother's view that deep coding is becoming less valuable as AI tools advance. The post reflects a genuine tension for CS students globally — whether foundational engineering rigor remains essential for top-tier roles and funded graduate programs, or whether AI proficiency now competes as a primary career differentiator. The student's stated goals (high GPA for funded Master's, FAANG employment, becoming a skilled engineer) hinge partly on this choice.

    The thread itself is a discussion post with no resolved conclusion or expert verdict stated in the body — the student posed the dilemma but did not report advice received or a decision made. The post was marked [D] (likely indicating 'discussion' tag on the subreddit).

  5. 5

    GPT-2's 32,070-token vocabulary mapped in interactive hyperbolic space

    A researcher has visualized GPT-2-small's complete vocabulary of 32,070 tokens inside a Poincaré ball (a mathematical representation of hyperbolic space) using the model's raw token embeddings, creating an interactive 3D explorer that runs on mobile devices. The vocabulary's underlying structure is forest-like—one large tree of about 2,300 tokens, several hundred smaller family trees, and around 6,700 isolated tokens—which fits naturally in hyperbolic geometry where space expands exponentially with distance, rather than in flat 2D space where trees distort.

    Users can interact with the visualization directly in a browser on their phone by dragging to rotate, pinching to zoom, and tapping tokens to center and explore neighboring relationships; this is a Möbius translation, the canonical navigation method in hyperbolic geometry.

What to Watch

Watch Alphabet's earnings report on July 22 for clues about Google Cloud's trajectory and when Gemini 3.5 Pro might reach users—momentum in cloud infrastructure and AI model releases will shape how quickly these tools move from labs into real-world applications. Meanwhile, code-fixing systems that combine planning, multiple agents, and sandboxing are proving they can solve real bugs without cutting corners, suggesting the next generation of AI coding assistants may be more reliable than today's shortcuts-prone versions.

Sources

Share this with a friend

Send today's roundup to anyone who wants to keep up.

Get daily AI news free with AIToday

200+ AI sources, summarized in 1 minute. Email / LINE / Slack.

Sign up free