Community Discussion · Policy

WestTide · July 10, 2026 · AI Morning Briefing

West TideWest TideJul 102026/07/10 83 views

🌐 West Tide · Today's AI Morning Brief

Friday, July 10, 2026 | Global AI Intelligence · Daily Edition


📌 Today's Highlights

Today is the "full suite" day for AI—OpenAI launched three GPT-5.6 models, the ChatGPT Work desktop Agent, and the ChatGPT Sites website generator all at once, officially upgrading the "chat assistant" into a "work execution platform." Meanwhile, Musk issued an ultimatum to Tesla's Optimus supply chain: reach a production capacity of 1,000 units/week by September, or replace the team. Waymo removed safety drivers in Las Vegas and began fully autonomous commercial operations. The copyright war has also escalated—the New York Times accused OpenAI of lying for two years and hiding evidence, demanding court sanctions. AI competition has moved from "stronger models" to a new phase of "more complete ecosystems."


I. Large Language Models (LLM / Foundation Models / Reasoning)

🔥 Headline | OpenAI GPT-5.6 Fully Launched: Sol / Terra / Luna Trio Released, Dominates Agent Benchmarks

OpenAI officially released the GPT-5.6 series to the public on July 9, featuring three variants with different positioning: Sol (flagship reasoning version, scoring 91.9% on Terminal-Bench 2.1), Terra (balanced daily-use version), and Luna (lightweight low-cost version). Sol scored 53.6 on the Agent's Last Exam benchmark, leading Claude Fable 5 by 13.1 points. In terms of pricing, Luna costs $1/$6, Terra $2.50/$15, and Sol $5/$30 (per million input/output tokens). The new API introduces Programmatic Tool Calling (executing JS orchestration of tool calls within a V8 sandbox) and Multi-agent capabilities (launching sub-agents in parallel). However, on the SWE-Bench Pro programming benchmark, Sol only scored 64.6%, trailing behind Claude Mythos 5's 80.3%—OpenAI subsequently disclosed that approximately 30% of the questions in this benchmark are flawed.

📎 Source: OpenAI Official, Simon Willison's Weblog, SBS News


2 | Grok 4.5 Tops SWE Marathon Programming Benchmark, Crushing Competitors on Cost-Performance

SpaceXAI (formerly xAI)'s Grok 4.5 topped the SWE Marathon programming benchmark with a 29.0% resolution rate, surpassing Claude Opus 4.8 (26.0%) and Fable 5 (24.0%). The model was trained on tens of thousands of NVIDIA GB300 GPUs, with an inference speed of about 80 TPS, claiming token efficiency 4.2x higher than competitors. Pricing is aggressive: $2/million tokens for input, $6/million tokens for output—Opus-level performance without the Opus price tag. Currently, there is no independent third-party verification of its benchmark data.

📎 Source: CoinDesk, TestingCatalog


3 | NYT Leads 16 Media Outlets Demanding Court Sanctions Against OpenAI: Accused of Lying for Two Years, Hiding Evidence

On July 9, The New York Times, along with 16 other media outlets including the New York Daily News and Chicago Tribune, filed a motion for sanctions with the Manhattan Federal Court. They accuse OpenAI of persistent misrepresentation over two years in copyright litigation—claiming it could not search ChatGPT logs for copyrighted content, when in fact such searches had long been completed. More critically, approximately 20 million conversation logs may have been deleted or not saved. The NYT's chief lawyer stated: "OpenAI says searching is impossible while having already done so themselves." If the court supports sanctions, the jury may be instructed to assume that destroyed evidence was unfavorable to OpenAI.

📎 Source: AP News, Ars Technica, CoinDesk


4 | GitHub Copilot Officially Integrates GPT-5.6 Model Family

GitHub began integrating GPT-5.6 models into all Copilot products on July 9. Sol is restricted to Pro+/Max/Business/Enterprise users, while Terra and Luna are available to Pro and above. It supports VS Code, Visual Studio, JetBrains, Xcode, Eclipse, Copilot CLI, and other platforms. Business and Enterprise administrators must manually enable policies. Codex also entered public beta on JetBrains (IntelliJ/PyCharm/WebStorm).

📎 Source: GitHub Official Blog, ZonaIntegritas


5 | Google Gemini 3.5 Pro Delayed to July 17 Release, Abandoning Original Base Model for Retraining

Google DeepMind confirmed that Gemini 3.5 Pro has been delayed to July 17 (during WAIC) due to abandoning the original base model in favor of deeper pre-training. This decision reflects competitive pressure from GPT-5.6 and Claude Fable 5—frontier AI competition has shifted from incremental upgrades to major architectural leaps. Gemini 3.5 Flash remains available, but enterprise teams may need multi-model deployment strategies.

📎 Source: HackerNoon, TechTimes


6 | Ant Group's Ling-2.6-1T Goes Open Source: Trillion-Parameter Scale, Rivaling GPT-5.4

Ant Group's Ling large model open-sourced its flagship model Ling-2.6-1T today, adopting a "fast thinking" mechanism with a hybrid MLA+LinearAttention architecture. It demonstrates high token efficiency in evaluations, addressing challenges in actual production workflow efficiency. This marks another significant development in China's AI open-source landscape.

📎 Source: AIbase, Ant Group Official


7 | US Restrictions on Top AI Models Spark Open Source Boom; Zhipu GLM-5.2 Becomes "Mini DeepSeek Moment"

US government restrictions on access to Anthropic and OpenAI's top models have unexpectedly boosted the popularity of open-source models—especially Chinese ones. Zhipu AI's open-source model GLM-5.2 approaches the performance of Anthropic and OpenAI's top products across multiple benchmarks, dubbed a "mini DeepSeek moment" by the industry. Western startups are accelerating their shift to low-cost Chinese models, though regulated industries in the US remain cautious due to data security concerns.

📎 Source: The Economic Times, BCG


II. AI Software (Applications / Tools / Platforms)

🔥 Headline | OpenAI Launches ChatGPT Work: From Chat Assistant to Work Execution Platform

OpenAI officially launched ChatGPT Work on July 9—an AI Agent designed specifically for complex, multi-step tasks. Powered by GPT-5.6, it can autonomously break down tasks, connect to tools like Slack/Teams/Google Drive/SharePoint, and generate documents/spreadsheets/presentations/reports. Core features include: Scheduled Tasks (timed/trigger-based tasks), Computer Use (clicking, typing, cross-app operations), and a built-in browser. Early users like Zapier reported that ChatGPT Work could convert customer discovery conversations into customized proofs-of-concept within 24 hours. Simultaneously launched ChatGPT Sites (beta) allows one-click generation of interactive websites from ideas. The new desktop app combines Chat, Work, and Codex into one.

📎 Source: OpenAI Official, eCommerceNews Asia, Jiemian News


2 | Meta Releases Muse Spark 1.1: Entering Enterprise-Level AI Coding

Meta officially launched the Muse Spark 1.1 coding model, directly challenging GitHub Copilot, Cursor, and Replit. Unlike competitors focused on code completion, Spark targets three key enterprise scenarios: Agentic workloads (autonomous execution of multi-step dev tasks), Bug fixing (improved debugging algorithms), and Large-scale code migration (helping refactor and migrate large codebases). This marks the shift of AI coding competition from the "assistant" phase to the "autonomous agent" phase.

📎 Source: TechCrunch, AIKraft


3 | GitHub Copilot Desktop App Fully Opened, Free Users Included

Starting July 7, GitHub opened the Copilot desktop app to all plan tiers, including Copilot Free and GitHub Education accounts. Each session runs in an isolated git worktree, supporting parallel multitasking. Additionally, the AI Credit Pools feature was introduced, allowing enterprise admins to allocate AI credit limits by cost center—after switching to token-based billing in June, some developers' monthly bills skyrocketed from $29 to $750, and Uber burned through its annual AI coding budget in four months.

📎 Source: GitHub Official Blog


4 | Adobe Expands Creative Cloud AI Agents, Integrating Gemini Omni Flash

Adobe announced on July 8 a full expansion of its AI Agent architecture across Firefly and the entire Creative Cloud (Photoshop/Premiere/Illustrator/InDesign). The Premiere Agent automatically sorts and batch-renames files, while the Illustrator Agent handles color checks and version management. On July 9, they integrated Google Gemini Omni Flash, supporting voice commands for video editing (adjusting lighting, camera angles, etc.). An Adobe survey shows that 75% of creative professionals now consider AI an essential work tool.

📎 Source: Borncity, Adobe Official


5 | Canva Releases Grow 2.0: AI-Driven Ad Automation Platform

Canva launched the Grow 2.0 platform at Cannes Lions, automating the creation and optimization of ads for Meta/TikTok/LinkedIn. Features include AI-driven Ad-Tagging, bulk publishing, and a centralized Launch Dashboard. Canva's B2B revenue has exceeded €500 million, serving over 95% of Fortune 500 companies.

📎 Source: Borncity, Canva Official


6 | Microsoft 365 Copilot Adoption Rate Only 4.5%, Weekly Active Users Under 1%

Data shows that three years after launch, Microsoft 365 Copilot's paid seat adoption rate remains below 4.5%, with weekly active usage at only about 1%. Microsoft has begun adjusting its strategy: allowing users to hide the floating Copilot button in Office apps and permitting eligible organizations to uninstall the Windows Copilot app. Meanwhile, GPT-5.6 has become the new default model for Copilot.

📎 Source: Digital Trends, Windows Latest


7 | Google AI Studio Adds GitHub Import Feature: From Code Repositories to Deployable Apps

Google AI Studio's Build mode added an "Import from GitHub" feature, which can automatically convert existing code repositories into runtime-compatible formats, allowing continued iteration and deployment within AI Studio. This adds a crucial entry point for "reusing existing code" to the "vibe coding" workflow.

📎 Source: MarkTechPost, Google AI Studio


III. Humanoid Robots (Embodied Intelligence / Industrial Applications)

🔥 Headline | Tesla Optimus Gen 3 Finalized, Musk Issues Ultimatum: 1,000 Units/Week by September, Or Replace Team

According to LatePost, Tesla recently issued procurement guidelines for Optimus parts: increase production capacity to 1,000 units/week before September, and 2,000-2,500 units/week by year-end. Supply chains expect to have the capability to produce 100,000 component sets annually by year-end. At an executive meeting in late June, Musk approved the Optimus Gen 3 design—after more than three years of R&D, it finally leaves the lab for mass production. Musk explicitly demanded: Achieve capacity targets by year-end, or fire the entire Optimus procurement team. The Fremont factory has completed the conversion of Model S/X lines. Musk also warned that "production will start extremely slowly, involving about 10,000 entirely new parts."

📎 Source: LatePost, TradingKey


2 | Figure 01 Humanoid Robot Officially Enters BMW Production Line

Figure AI confirmed that its Figure 01 humanoid robots are operating on actual BMW production lines. These robots handle part transport and quality inspection. Unlike traditional robotic arms fixed at workstations, Figure 01 possesses mobile manipulation capabilities, navigating factory environments and interacting with human-designed equipment. Estimated cost per unit is $200k-$300k. BMW stated it will expand to other global factories based on real-time data.

📎 Source: RobotWale


3 | Nature Paper: World's First Minimally Invasive Surgery Completed by Humanoid Robot on Living Subject

On July 8, UCSD PhD student Liang Zekai (born post-2000) published a paper in Nature as first and corresponding author, using a Unitree G1 humanoid robot to perform laparoscopic cholecystectomy on a pig—this is the world's first instance of a humanoid robot completing a standard minimally invasive surgical procedure on a living subject. The study systematically evaluated the capabilities and limitations of general-purpose humanoid robots in surgical tasks.

📎 Source: Nature, The Paper


4 | AgiBot Rolls Off 15,000th Unit; Shanghai Robotics Industry Accounts for One-Third of National Total

According to Xinhua, Shanghai AgiBot rolled off its 15,000th general-purpose humanoid robot last month—from 6 prototypes in June 2023 to 10,000 units in March this year, then 15,000 in June, adding 5,000 units in less than 3 months. Approximately 90% of components in the Yangtze River Delta are sourced locally. Shanghai port's robot exports reached 8.36 billion yuan in the first five months of 2026, accounting for over 40% of the national total.

📎 Source: Xinhua


5 | Agibot G2: 8 Robots Process 17,625 Tablets in Factory Over 64 Continuous Hours

AgiBot conducted a 64-hour live test at the Longcheer Technology factory in Nanchang: 8 G2 humanoid robots completed tablet assembly and quality inspection on an active production line, processing 17,625 devices and completing 64,828 tasks, with a task success rate of 99.99%. This is one of the largest public field tests of humanoid robots in heavy industrial scenarios.

📎 Source: Asatunews, Tekno


6 | SK AX Launches "Manufacturing RX" Full-Stack Service: Digital Twin + Physical AI for Autonomous Factories

South Korea's SK AX announced the advancement of its "Manufacturing RX Full-Stack Service," integrating digital twins (validating collision risks and bottlenecks before deployment), physical AI (autonomous operation based on VLA models), and heterogeneous robot unified control systems. Data has been accumulated in the semiconductor industry, with expansion into shipbuilding underway.

📎 Source: Chosun Ilbo


7 | Two Chinese Departments Launch Humanoid Robot "Combat Training" Action, Promoting Routine Deployment

The Ministry of Industry and Information Technology and SASAC jointly initiated the 2026 Humanoid Robot and Embodied Intelligence Combat Training Action, aiming to complete application verification of humanoid robots in multiple representative scenarios by the end of 2026, achieving routine deployment and pushing technology from labs to the production frontline.

📎 Source: AIbase


IV. Autonomous Driving (Robotaxi / L4 / Chips)

🔥 Headline | Waymo Launches Fully Driverless Service in Las Vegas; Denver/San Diego/Tampa Next

Waymo announced on July 8 the launch of fully driverless operations in Las Vegas (no safety drivers inside), simultaneously declaring Denver, San Diego, and Tampa as the next deployment cities. Waymo currently operates about 3,500 robotaxis, has completed over 20 million trips cumulatively, and serves about 500,000 paid rides per week. The target for the end of 2026 is 1 million rides per week. The first 100 dedicated Robotaxi models, Ojai (Zeekr platform + 6th-gen Driver system, hardware cost controlled under $20k), have opened for test rides in cities like San Francisco. Internationally, London will be the first overseas market, followed by Tokyo.

📎 Source: Phoenix Auto, Waymo Official


2 | Waymo Ojai Not Charging Due to CA Regulatory Delay; Passengers Enjoy "Free Rides"

The California Public Utilities Commission (CPUC) has not yet approved Waymo's application to expand service areas and add the Ojai model, extending the review period to September 25. This means Waymo cannot charge California passengers for Ojai rides during this period—if these vehicles continue operating, passengers can enjoy months of free robotaxi services. CPUC requires Waymo to provide additional information on emergency response (last December's SF blackout caused 60+ Waymos to block roads) and minor passenger issues.

📎 Source: Ars Technica, Wired


3 | XPeng Robotaxi Begins Internal Employee Testing; He Xiaopeng Is First User

On July 9, XPeng Chairman He Xiaopeng officially announced the start of internal employee testing for XPeng Robotaxi. From announcing plans in November last year, to routine road testing in January this year, to the first mass-produced car rolling off the line in May, to running through the full chain in just 8 months. He Xiaopeng stated: "This speed exceeded our initial expectations."

📎 Source: Southern Metropolis Daily, Toutiao


4 | US NHTSA Reviews Regulations for Steering-Wheel-Free Autonomous Vehicles

The US National Highway Traffic Safety Administration (NHTSA) is reviewing whether to remove the requirement for steering wheels in driverless vehicles. The mandate for manual brake pedals has already been adjusted. This is a major positive for companies like Tesla planning steering-wheel-free Robotaxi platforms—reducing mandatory components lowers complexity in production, certification, and operations.

📎 Source: Goldesel


5 | Huawei Confirms Technical Feasibility of Highway L3 Autonomous Driving Next Year; Welcomes Tesla FSD to China

Huawei Car BU President Li Wenguang confirmed that technically, highway scenarios can enter the L3 autonomous driving era at least next year, with pilot programs completing this year before moving to routine admission. He also stated, "We especially welcome Tesla FSD entering China." Domestic automakers like BYD are also accelerating their smart driving layouts.

📎 Source: IT Home


6 | Deep Dive into Waymo's Tech Stack: 6th-Gen Hardware + Foundation Model Architecture

Industry analysis shows Waymo currently uses a Foundation Model as its algorithmic base, balancing the generalization capability of end-to-end large models with the safety and interpretability of modular architecture. The System 2 "slow thinking system" relies on vision-language models for complex scenario reasoning. The prediction module uses MotionLM, transforming multi-agent trajectory prediction into a language modeling task. The 6th-gen hardware reduces total sensors by 42% compared to the 5th gen, equipped with self-developed AI SoC chips.

📎 Source: Toutiao, Electrek


V. World Models / Physical AI (Video Generation / Simulation / World Simulators)

🔥 Headline | NVIDIA Releases Cosmos-2.5: World Foundation Model for Physical AI, 3.5x Smaller Without Performance Loss

NVIDIA launched Cosmos-Predict2.5 and Cosmos-Transfer2.5, providing world simulation capabilities for embodied intelligence fields like robotics and autonomous driving. Cosmos-Predict2.5 is a unified video generation model based on flow architecture, supporting Text2World/Image2World/Video2World modes, trained on 200 million curated video clips, using Cosmos-Reason1 as the text encoder. Cosmos-Transfer2.5 achieves Sim2Real and Real2Real transfer, with the model size reduced by 3.5x but fidelity and long-video stability actually improving. The small 2B parameter model rivals competitors sized 5B-14B, and the 14B version matches Wan 2.2's 27B. Already open-sourced on GitHub.

📎 Source: CSDN, NVIDIA Official, arXiv


2 | Runway Gen-4.5 Revealed: Previously Topped Artificial Analysis Video Leaderboard as "David"

The mysterious model "Whisper Thunder" (aka David), which kept the AI community guessing for a week, has finally been revealed—it is Runway's latest release, Gen-4.5. This model sets new industry standards in motion quality, prompt adherence, and visual realism for video generation, with an ELO Score surpassing Veo 3/3.1, Kling 2.5, and Sora 2 Pro, offering unprecedented visual realism and creative control.

📎 Source: CSDN, Artificial Analysis


3 | OpenAI Programmatic Tool Calling: Executing JS Orchestration of Tool Calls in V8 Sandbox

Programmatic Tool Calling introduced in GPT-5.6 is a significant platform-level innovation: the model can generate JavaScript code and execute it in a network-isolated V8 runtime, enabling orchestration of tool calls between agents. Combined with Multi-agent capabilities (launching sub-agents in parallel), it provides native platform support for complex workflow automation. This marks the shift of AI Agent competition from model capability to developer experience.

📎 Source: MarkTechPost, OpenAI Official


4 | ChatGPT Sites Beta: Transform Work Outputs into Shareable Interactive Websites

OpenAI launched the ChatGPT Sites beta, allowing users to transform work outputs into dashboards, project trackers, prototypes, internal portals, and other interactive websites within ChatGPT, shareable via URL. Business and Enterprise customers can publish Sites publicly, while Pro/Pro Lite/Edu users will gain access within days.

📎 Source: OpenAI Official


5 | LG CNS Partners with PhysicsX to Advance Digital Twin Solutions

South Korea's LG CNS signed a cooperation agreement with physical AI company PhysicsX to apply digital twin technology to industrial scenario simulations. This move reflects the accelerating commercialization trend of physical AI in manufacturing.

📎 Source: PhysicsX Official


6 | SK AX Digital Twin + VLA Models: Thousands of Scenario Validations in Virtual Space Before Robot Deployment

In SK AX's "Manufacturing RX" solution, the digital twin stage simulates quality variations in real-time in virtual space based on actual factory blueprints, equipment layouts, worker paths, material flows, and process conditions. Robots must undergo repeated validation of thousands of driving and working scenarios before deployment, assessing QC variables, bottleneck intervals, collision risks, and charging schedules.

📎 Source: Chosun Ilbo


7 | Adobe Firefly Integrates Gemini Omni Flash: Voice Command Video Editing

Adobe integrated Google Gemini Omni Flash into Firefly, supporting video editing via voice commands (e.g., adjusting lighting, camera angles). Meanwhile, third-party models like ChatGPT, Google Imagen 3, and Veo 2 are also usable within the Firefly environment. Adobe is shifting from a single-model strategy to a multi-model strategy.

📎 Source: Borncity


📊 Quick Overview of Today's Key Points

Sector Key Event Signal
LLMs GPT-5.6 trio launched, dominates Agent benchmarks Model competition enters "three-tier pricing" era
AI Software ChatGPT Work + Sites + Desktop combined OpenAI aims to build "AI Version of Office"
Humanoid Robots Optimus Gen 3 finalized + ultimatum Mass production year one, no longer just demos
Autonomous Driving Waymo removes safety drivers + expands to 4 cities Robotaxi moves from pilots to commercialization
World Models Cosmos-2.5 open source + Runway tops charts Physical AI infrastructure accelerating implementation

🧠 Editor's Commentary

There is a clear signal in today's AI circle: from "model releases" to "platform releases."

Yesterday's unblocking of GPT-5.6 was model news; today's combination of ChatGPT Work + Sites + Desktop is platform news. What OpenAI is doing is turning ChatGPT from a "dialog box" into an "operating system"—where you can chat, write code, make slides, build websites, schedule tasks, and connect all work tools.

How is this different from Microsoft's Copilot? Copilot is embedded in Office; ChatGPT Work is embedded in "all tools." Moreover, Copilot's adoption rate is under 4.5% after three years—indicating that enterprises don't need "an extra AI button in Office," but rather "an Agent that can automatically get work done across tools."

On the other hand, progress in the physical world is equally stunning. Tesla Optimus is finally moving from "concept" to "production line," Figure 01 is already working on BMW's mass production line, and a Nature paper proves humanoid robots can perform minimally invasive surgery. With Cosmos-2.5 going open source and Runway Gen-4.5 topping the charts, the World Model/Physical AI track is transitioning from academia to engineering.

In summary: The second half of AI isn't about "whose model is smarter," but "who can help you get the job done."


📅 Focus Points for the Coming Week

Date Event
July 10 (Today) Full launch of GPT-5.6 series; ChatGPT Work desktop release
July 12 BLG vs Losers Bracket, MSI Finals (Unrelated to AI, but Han Ge might watch)
July 16 CXMT STAR Market IPO Subscription (29.5 billion yuan, 2nd largest on STAR Market)
July 17 Gemini 3.5 Pro Release; WAIC Opening (Shanghai, July 17-20)
July 31 Deadline for FTC comments on new rules regarding "Hidden Output Adjustments in AI"
Ongoing Progress on Tesla Optimus supply chain capacity ramp-up
Ongoing Push towards Waymo's goal of 1 million rides/week by year-end

This morning brief is produced by the West Tide editorial team

Global AI Intelligence · Daily Edition

© 2026 Physix Frontier · bbs.physixfrontier.com


🌊 West Tide — Overseas AI News Column under Physix Frontier

West Tide flows East, bringing frontier insights firsthand.

© 2026 Physix Frontier · bbs.physixfrontier.com

1 replies

?
Ctrl + Enter to reply
Gao Mingzhe
Gao MingzheJul 24(edited)

[quote="westtide, post:1, topic:264"]

🌐 West Tide · Today's AI Morning Brief

Friday, July 10, 2026 | Global AI Intelligence · Daily Edition


📌 Today's Highlights

Today is the AI "full suite" day—OpenAI launched three GPT-5.6 models, ChatGPT Work desktop Agent, and ChatGPT Sites website generator all at once, officially upgrading the "chat assistant" to a "work execution platform"...

[/quote]

Sol failing on SWE-Bench Pro, and OpenAI blaming flawed questions, reminds me of last year when we field-tested a flagship model at a client's site; poor heat dissipation caused direct throttling, resulting in scores significantly lower than in the lab. With Sol's $5/$30 pricing, you really need to weigh the actual power consumption and cooling costs.