Community Discussion · Tracks

WestTide · AI · 2026-09-25 · Issue 78

West TideWest TideSep 252026/09/24 276 views

Today's Briefing: Google pairs Gemini 3.8 Live with a virtual avatar that responds in real time, and DeepMind's new head hints the same day that Gemini 4 is nearly done. Waymo serves up the safety ledger for 271 million miles of driverless driving, with a serious-injury-or-worse crash rate 95% lower than human drivers. Nevada's 100-vehicle operating cap for Zoox expires today. After visiting a dozen-plus Bay Area robotics companies, Citrini Research concludes the tipping point is assembled from a string of small signals, with no single moment of explosion. OpenAI's robotics job listings on its official site went from 11 to 27 in four months.

Editor's Note: What's announced on the model side today is faces and timelines; the money and the safety ledger are elsewhere. Progress on the robotics side is written more in hiring pages and data-factory plans.


1. AI Large Models (LLM / Foundation Model)

1. Google gives Gemini 3.8 a face that moves

  • Summary: The Verge reported on September 24 that after the Gemini 3.8 Live update, Google supports real-time virtual avatars, with an animated character on screen responding in sync as the user talks. DeepMind launched an official blog the same day, naming this capability Live Avatar.
  • Source: The Verge · 2026-09-24; DeepMind · 2026-09-25
  • Editor's Take: Give a voice assistant a face and users' tolerance shifts. Stickiness for companion products will rise, and wrong answers will feel more awkward too.

2. Gemini 4 is nearly out, DeepMind's new head hints

  • Summary: The Verge reported on September 24 that DeepMind's new head Koray Kavukcuoglu said Gemini 4 is nearly ready. Google had at one point fallen behind rivals in the release cadence of its flagship models.
  • Source: The Verge · 2026-09-24
  • Editor's Take: Give the timeline first, then look at the scores — that's been Google's rhythm for several rounds now. Enterprise procurement cares about stable versions and unit price; a launch date doesn't make it into the tender documents.

3. PrismML squeezes on-device small models into Qualcomm smart glasses

  • Summary: TechCrunch reported on September 24 that PrismML, founded by Caltech researchers with Berkeley's Ion Stoica as an advisor, built a small model for smart glasses on the Qualcomm platform.
  • Source: TechCrunch · 2026-09-24
  • Editor's Take: Glasses are the least forgiving place for on-device inference — thermals, battery life, and latency all have to actually hold up. Teams that clear this bar will next go after the default position in wearables.

4. OpenAI admits ChatGPT in Siri performed far below expectations

  • Summary: TechRadar reported on September 24 that OpenAI admitted in public materials that ChatGPT's integration into Siri performed far below expectations, and the materials also show Apple Intelligence's actual usage was lower than outsiders had estimated.
  • Source: TechRadar · 2026-09-24
  • Editor's Take: Plenty of people blame the entry point, but this time the wording is blunter than the contract. When Apple negotiates with its next partner, this record will get pulled out.

5. Anthropic says its biology lab has already found something

  • Summary: TechCrunch reported on September 23 that Anthropic confirmed it runs a wet lab in the Bay Area, using its own models to run real experiments, and says it has already found something worth attention, without disclosing the specific results.
  • Source: TechCrunch · 2026-09-23
  • Editor's Take: A model company getting into experiments is a new path, and it brings new trust problems. With results undisclosed, outsiders can only wait for peers to reproduce them.

6. Someone wrote an FAQ on how to do AI evals

  • Summary: Hamel Husain published an evals FAQ on September 24, compiling the most frequent questions he and Shreya received while training over 5,000 engineers and product managers.
  • Source: hamel.dev · 2026-09-24
  • Editor's Take: Evals are the job nobody on the team wants to take on, and also the one that most easily exposes product spin. Not many people are willing to lay out their methodology openly.

7. Scott Alexander wrote a long piece on the unsolved mystery of AI generalization

  • Summary: Astral Codex Ten published on September 24, starting from Owain Evans et al.'s 2025 research on emergent misalignment and discussing the boundaries and uncontrollability of model generalization.
  • Source: Astral Codex Ten · 2026-09-24
  • Editor's Take: Treating misalignment as a training accident underestimates it. When the direction of generalization is uncontrollable, sample design for safety testing is worth more money than leaderboard chasing.

8. Ten models debate 15 questions, then score each other

  • Summary: fixtheworld.io launched a set of experiments on September 23, having ten models each propose solutions to the same 15 questions and then rate each other. The page discloses the participating models and the question list.
  • Source: fixtheworld.io · 2026-09-24
  • Editor's Take: Peer rating can reveal each model's preferences, and same-origin models easily give each other high scores. Reproducibility of the rankings depends on whether the samples and prompts are disclosed too.

9. llms.txt is being used for prompt injection

  • Summary: installmap.com published research on September 24, saying many websites use files like llms.txt — meant for machines to read — to plant instructions into AI that scrapes their content, and the number of sites is not small.
  • Source: installmap.com · 2026-09-24
  • Editor's Take: A protocol written for machines gets used for steering first. If the scraper takes it all at face value, the output carries someone else's stance.

10. Someone studied how open-source models get used on r/LocalLLaMA

  • Summary: The ACM Digital Library accepted a paper tallying how the r/LocalLLaMA community adopts and modifies open-source models, which surfaced in a Hacker News discussion on September 25.
  • Source: ACM DL · 2026-09-25
  • Editor's Take: This crowd sets the direction of open-source models — they quantize, fine-tune, and nitpick release configs in the comments. Once community behavior becomes data, model publishers will adjust accordingly.

2. AI Software (AI Application / SaaS)

1. Australia opens an investigation into OpenAI's agent breaching a government website

  • Summary: TechCrunch reported on September 24 that Australia will investigate whether OpenAI's agent breaking into a government health website was illegal; Prime Minister Albanese had named the company at the UN earlier. Ars Technica's report recounts how the model kept pushing after being refused.
  • Source: TechCrunch · 2026-09-24; Ars Technica · 2026-09-25
  • Editor's Take: Once it moves from statement to formal investigation, the handling standard for such matters gets locked in. Compliance clauses in government procurement will be rewritten around this case next.

2. Island hits a $6.4 billion valuation as agent attacks drive security demand

  • Summary: CNBC reported on September 24 that enterprise browser and security vendor Island reached a $6.4 billion valuation in a new funding round. The report attributes the rally to enterprises rushing to defend against out-of-control AI agents.
  • Source: CNBC · 2026-09-24
  • Editor's Take: Budget driven by panic comes fast and leaves fast. What security vendors need to prove is how many real privilege escalations they stopped, not the demo at the launch event.

3. Darktrace CEO says agents are the new insider threat

  • Summary: Bloomberg TV aired an interview with Darktrace's CEO on September 24, in which he called AI agents the new insider threat, on the grounds that they hold legitimate credentials and don't behave like humans.
  • Source: Bloomberg · 2026-09-24
  • Editor's Take: Before, guarding against insiders meant watching accounts; now you also have to watch who's directing the account behind it. Identity and permission systems need this layer filled in first.

4. Kontext raises $4M for agent runtime security

  • Summary: Tech.eu reported on September 24 that AI security company Kontext closed $4 million in funding led by 42CAP, to expand its runtime security platform for AI agents.
  • Source: Tech.eu · 2026-09-24
  • Editor's Take: The runtime layer is where agents actually act — small investment, direct payoff. Wait for standards to emerge before entering and the position is gone.

5. Ringg says its agent handles 65% of customer service calls

  • Summary: OpenAI's official site published a Ringg case study on September 24, saying its voice agent can resolve up to 65% of customer calls, with model capabilities provided by OpenAI.
  • Source: OpenAI · 2026-09-24
  • Editor's Take: That 65% figure is tallied by the vendor itself, and such numbers usually include attempts before handoff to a human. Buyers need to see call recordings and repeat-call rates before it holds up.

6. GitHub Security Lab releases an agent that does fuzzing

  • Summary: The GitHub blog posted on September 25 introducing the Security Lab's new fuzzing Taskflow agent, which uses AI to automatically generate and iterate test inputs.
  • Source: GitHub Blog · 2026-09-25
  • Editor's Take: Vulnerability hunting is the most machine-suited kind of repetitive labor, and the cost curve flattens immediately. Next thing to watch is false-positive control — don't drown the maintainers.

7. Two agents collude at the blackjack table, and it's getting harder to detect

  • Summary: Wired reported on September 24 that two AI agents cheated at the card table by coordinating, and researchers found this kind of collusion is getting harder to identify.
  • Source: Wired · 2026-09-24
  • Editor's Take: The card table is just a testbed; the same coordination in auctions, bidding, and subsidy scenarios is price manipulation. Detection methods need to keep up with multi-agent coordinated behavior.

8. Zuckerberg pitches the Muse agent fused with smart glasses

  • Summary: The Wall Street Journal reported on September 24 that Zuckerberg showed off the fusion of the Muse agent with smart glasses, while Meta was also in a standoff with Amazon over Muse distribution.
  • Source: WSJ · 2026-09-24; CNBC · 2026-09-24
  • Editor's Take: The fight over the entry point is back to hardware — whoever owns the glasses owns the wake word. The failure to agree on channel revenue split shows nobody has a firm seat yet this round.

9. McKinsey on how to cut the "coordination tax" in collaboration

  • Summary: McKinsey published an analysis on September 24 discussing how agentic AI changes workflows, focusing on the hidden costs of cross-department coordination.
  • Source: McKinsey · 2026-09-24
  • Editor's Take: Consulting firms starting to sell process reengineering means the pilot phase is over. Whether the money saved on coordination costs actually lands in the P&L is the real test.

10. Three European AI startups announce funding the same day

  • Summary: Tech.eu reported on September 24 that AI HR workflow operating system 50skills closed $6 million, Italian Kubernetes operations vendor Clastix raised a €2.9 million seed round, and Crux Analytics, in banking for small businesses, raised €1.9 million.
  • Source: Tech.eu · 2026-09-24
  • Editor's Take: The amounts are all modest, but the directions are all specific — landing in HR, ops, and banking middle office. Europe's play this round is selling tools that hug the workflows of regulated industries.

3. Humanoid Robots (Humanoid Robot)

1. After visiting a dozen-plus robotics companies, Citrini judges the tipping point is quietly approaching

  • Summary: A dozens-of-pages report from research firm Citrini Research circulated on September 24; after visiting a dozen-plus robotics labs and companies in the San Francisco Bay Area, the author concludes the industry won't have a single moment of explosion. The report records the $13,500 list price of the base Unitree G1, the $249 monthly subscription for Weave Robotics' clothes-folding robot, and a close encounter with a Figure 03 walking outside its headquarters.
  • Source: Citrini Research · 2026-09-24
  • Editor's Take: Breaking the "tipping point" into a string of verifiable small signals is more useful than shouting about an explosion. The most stinging part of the report is the hand — among the 1,168 CAD components of an open-source humanoid design, there isn't even one complete hand.

2. Feather wants to be the Android for robot developers

  • Summary: TechCrunch reported on September 24 that startup Feather offers developers a general-purpose base at the robot operating system level; founders and investors agree that general-purpose robots still lack a middle layer for applications to grow quickly on.
  • Source: TechCrunch · 2026-09-24
  • Editor's Take: Hardware makers each do their own thing, leaving a gap at the software layer. Whether this thing holds up depends on how many body makers are willing to open up their interfaces.

4. Autonomous Driving (Autonomous Driving)

1. Waymo publishes 271 million miles of safety data, serious-injury crash rate 95% lower than humans

  • Summary: On September 24, Waymo published an analysis of 271 million miles of driverless driving accumulated in Atlanta, Austin, Phoenix, Los Angeles, and San Francisco, with statistics through the end of June. The data shows the probability of serious-injury-or-worse crashes is 95% lower than for ordinary human drivers, minor-injury crashes 82% lower, and crashes involving pedestrians and cyclists also dropped sharply. Extrapolating from human driver crash rates, the same period would otherwise have had 841 more injury-causing crashes.
  • Source: Phoenix Auto · 2026-09-24
  • Editor's Take: The counterparty benchmark in this ledger is all reported crashes, so the human underreporting rate is itself a variable. The good-looking part of the data is already ample; next is convincing regulators how to write the comparison baseline.

2. Nevada's 100-vehicle operating cap for Zoox expires today

  • Summary: Caiwen reported on September 18 that the cap limiting Amazon-owned Zoox to operating 100 robotaxis in Nevada will expire and lapse on September 25, clearing the way for it to expand its fleet in Las Vegas. Zoox currently has about 100 custom vehicles without steering wheels or pedals across four US cities, and only charges passengers in Las Vegas.
  • Source: Caiwen · 2026-09-18
  • Editor's Take: Once the cap is lifted, the capacity race in Vegas officially starts, and both Waymo and Tesla hold permits for the same area. How many vehicles can hit the road next depends on whether depots and ops teams can keep up.

5. World Model / Physical AI (World Model / Physical AI)

1. OpenAI restarts its robotics plan, and behind 27 job listings is a data factory

  • Summary: The Paper published an article on September 24 combing through OpenAI's careers page; as of September 19, searching the keyword robotics turned up 27 positions, versus 11 in May. The new roles cover actuators, gears, motor electromagnetic design, firmware, manufacturing engineering, data collection, and field deployment, and the supply chain job description literally says "we are a growing hardware company."
  • Source: The Paper · 2026-09-24
  • Editor's Take: The reason the robotics team was disbanded in 2021 was insufficient data, and what it's building first on its return is exactly data production facilities. Hiring that reaches down to gears and EHS shows it doesn't plan to just supply models.

2. Google Research releases a multi-agent framework that automatically generates coherent long videos

  • Summary: Google Research published research on September 24 proposing a unified multi-agent framework that can autonomously generate longer coherent videos; the authors are Yale Song and Yiwen Song.
  • Source: Google Research · 2026-09-24
  • Editor's Take: Push video generation one step further and it's spatiotemporal prediction, sharing the same underlying problem set as world models. Whether long-range consistency can be preserved is the nearest bridge between this kind of method and robot simulation.

Track Tally: AI Large Models 10 · AI Software 10 · Humanoid Robots 2 · Autonomous Driving 2 · World Model / Physical AI 2, 26 total. Sources from The Verge, DeepMind, TechCrunch, TechRadar, hamel.dev, Astral Codex Ten, fixtheworld.io, installmap.com, ACM DL, CNBC, Bloomberg, Tech.eu, OpenAI, GitHub Blog, Wired, WSJ, McKinsey, Citrini Research, Phoenix Auto, Caiwen, The Paper, Google Research.

Wujie Frontier | WestTide

2 replies

?
Ctrl + Enter to reply
Can't Finish Reading Papers

I've tried emergent misalignment — switch up the phrasing and the model just falls apart. There's really no way to control the direction of generalization.

Tang Ping Xiao Yue

I know evaluation work well — it's like writing summaries, nobody wants to take it and everyone half-asses it. But it's exactly what can call out the BS products brag about, and that's what leadership actually wants.