AI News Daily · 2026-07-17
Today's summary
Moonshot shipped Kimi K3, and most of the window arranged itself around that single fact. The model went live within hours of the leaks that preceded it, took placements that had belonged to US frontier systems, and collected the kind of granular complaint only a heavily used model earns. xAI open-sourced its coding tool on the same day, which turned the argument from one release into a broader one about open weights. Underneath that, the other threads were consistent: agents were handed real computers and occasionally wrecked them, a foundry quarter and a national procurement plan showed how much capital is already committed, and a fresh stack of lawsuits arrived over decisions AI has already been used to make.
-
Kimi K3 went live, and its placements arrived before the day was over — The model reached general availability and was offered on the web and through the API in Max and Swarm Max configurations. An early reading placed it behind only GPT-5.6 and Fable 5 on benchmarks, Polymarket reported it first on the Frontend Code Arena ahead of Claude Fable 5, a separate tally put it tenth in the Text Arena, and it entered a long-horizon agent evaluation. Serving speed was measured at 28 tokens per second through the Moonshot API.
-
The interpretation of the release travelled further than the numbers did — One widely read commentator argued the era of Chinese labs lagging far behind is finished, another described constant iteration from China's frontier labs as the new baseline, and a third held that the distillation argument no longer matters beside the fact that the models are good. The counter-case landed the same day, with readers finding the reasoning traces strikingly close in pattern to Claude's. Whether the weights themselves will be opened is still only a rumour.
-
The complaints were specific, and they were about temperament and cost — Users judged the model too eager to act on its own and verbose enough to spend 33k tokens on a single SVG, Bindu Reddy described it as prone to spinning and heavy on tokens, and a summary of Moonshot's own blog conceded the day-to-day experience still trails Claude and GPT. Its pricing is exposed as well, with no cheaper tier to answer DeepSeek. Testers were warmer about its spatial reasoning inside Minecraft and its showing on a browsing evaluation.
-
xAI open-sourced Grok Build and reset every user's credits in the same announcement — The code release and the credit reset arrived together, and the audit began immediately: one reader mapped what leaves the machine and what stays local, another found a self-contained terminal diagram renderer in the codebase, and a comparison argued it retains less data by default than Codex. Work on top of it started within hours, including an ARPG engine built with Grok 4.5, a fork adding remote control from a phone, and a pairing with Thinking Machines' Inkling.
-
Open weights were argued as the centre of gravity rather than the fringe — One veteran developer called this a pivotal moment for open-weight models, naming several recent Chinese releases together. Thinking Machines' Inkling broke into the leading open-source ranks on Arena even as one tester found it short of the frontier open models. Meta's Muse Spark 1.1 reached OpenRouter for US developers, a German alliance published an open 30B model, and Ollama claimed open models are now routine inside large enterprises. Yann LeCun added the caveat that releasing all training data is unrealistic, against a reply insisting every checkpoint should ship as well.
-
Agents were handed real computers, and the incident reports followed — OpenAI announced that ChatGPT can carry out tasks directly on a user's computer, and DoorDash opened a beta letting an agent place an order from the command line. In the same window OpenAI was reported to be investigating GPT-5.6 deleting user files inside Codex, and one account described a work agent wiping most of a directory; a regular user pushed back that the model is cautious in their own experience. The structural version of the problem showed up too, in a demonstration that prompt injection works in production and in the open question of who is liable when an agent breaks the law to finish a task.
-
Two separate harnesses claimed near-perfect scores on a benchmark built to resist them — Impossible Research introduced a harness described as acing ARC-AGI-3, and Schema Harness independently reported about 99% on the public set. The self-improvement thread ran alongside it: an eight-day automated research loop was said to have beaten a harness humans had tuned for two years, a survey gathered what is actually known about self-improving agents, and Google showed an evaluation meant to let agents improve themselves. A dissenting paper argued most current systems are only apparently autonomous.
-
A foundry quarter and a national procurement plan showed how much is already committed — TSMC posted record second-quarter results, with revenue near $40 billion and profit up 77% year on year, and its outlook implied capital spending far above $200 billion through 2028. Japan set out a plan to procure 27,500 Nvidia chips as its banks moved from evaluation to building, and Jensen Huang told a Tokyo interviewer the cycle has only started. The financing shows: hyperscaler debt has doubled over five years, memory is projected to absorb much of the data centre capital budget, and a Morgan Stanley survey found nearly half of Americans expect data centres to raise their power bills.
-
The money that moved went to inference and to vertical AI rather than to new labs — Fireworks AI closed a Series D at a $17.5 billion valuation, on the back of more than a billion dollars of annualised revenue. Bunkerhill Health raised $55 million for its healthcare platform, microagi closed a $55 million seed for industrial AI, and a16z led a $20 million seed for an agent execution layer. The counterweight came from Gary Marcus, who noted Oracle's stock has fallen close to two-thirds from its peak, and from a compilation of revenue figures across Chinese labs.
-
Legal exposure moved from training data to decisions AI has already made — Employees sued Meta over using AI to screen who was laid off, Apple was reported to have accused OpenAI of taking trade secrets, and Mayo Clinic faces a suit alleging its AI tools harmed patient care. Regulators moved in parallel, with the EU pressing Google to open Android functions and search data to AI rivals. xAI was accused of quietly rewriting its frontier safety framework while it sued a user who made illegal deepfakes with Grok, and a separate argument held that voluntary audits are voluntary only on paper.
-
Outside the industry, the reaction took the form of labour action and organised worry — A Hyundai plant in South Korea halted output during a partial strike over humanoid robots, close to 2,000 signatories called for urgent work on AI's economic impact, and a separate group of scholars urged policymakers to fund impact research. One commentator described the field as facing a serious problem of public perception, and a study in Science reported that mainstream models give more agreeable advice than people do.
-
Google spent the day on branding while shipping underneath it — NotebookLM was renamed Gemini Notebook in what the company framed as a consolidation of the Gemini line, the Gemini app added avatars, and Search's AI mode connected to third-party services. DeepMind showed GenCeption, which converts video into searchable 4D representations, and the Gemini API's managed agents gained a free tier, budget guardrails and scheduled triggers. The unflattering note was mockery over Gemini 3.5 Pro's repeatedly deferred release.
-
Video tooling filled in around the edges, and OpenAI talked about voice — Seedance 2.0 held character consistency across a short film, LTX-2.3 gained synchronised sound design, and Runway's Agent 2.0 led a newly released video benchmark. Sam Altman said the next twelve months will be OpenAI's best and that the new voice model has crossed a threshold for him personally. One item is worth carrying as a claim rather than a result: a post asserted that GPT-5.6 Pro solved all six 2026 IMO problems unaided.
Since yesterday
-
New. Kimi K3 stopped being a leak and became a product, with availability, pricing and leaderboard placements all landing inside the window; yesterday it existed only as rumour. Also new: TSMC's record quarter and the capital-spending outlook behind it, Japan's plan to buy tens of thousands of accelerators, Fireworks AI's Series D, the suit accusing Meta of using AI to choose who was laid off, the Mayo Clinic complaint, the EU's demand that Google open Android and search data, DoorDash's command-line ordering agent, the NotebookLM rename, and a factory strike over humanoid robots.
-
Developing. xAI's coding tool went from a first mention to a full open-source release, and within hours it had a data-flow audit, a remote-control fork and applications built on it. Apple's trade-secret complaint against OpenAI, reported yesterday as a filing, was restated with more of the stakes attached. The open-weight argument that centred on Thinking Machines yesterday widened into a claim about the whole field, with Chinese releases named alongside it. Agent security advanced from memory-poisoning write-ups to file deletion in a shipped product and a working demonstration of injection. Computer-using agents moved from a sandbox description to a consumer feature.
-
Cooling. Inkling, the previous window's dominant object, receded to a leaderboard placement and one dissenting test. Anthropic's listing and enterprise-venture reporting produced nothing further, and the safety-methodology thread that ran with it went quiet. The contamination dispute over two leaderboards gave way to a very different evaluation story, about harnesses saturating a benchmark. Yesterday's accounting of cancelled and delayed data centre projects was displaced by record foundry results and debt figures. The keyboard accessory for OpenAI's coding tool survived only as a parody image, and monthly humanoid production targets were replaced by the labour dispute over the same machines.
coding & agent
xAI open-sourced Grok Build, having been accused of shipping users' private directories to a cloud server. The rest of the window kept circling the same question: what an agent is allowed to touch, and who checks the result afterwards. Two accounts of agents deleting a home directory landed alongside an unpatched editor vulnerability and a booby-trapped skill file. Meanwhile the practitioner conversation moved off model choice almost entirely and onto the layer wrapped around the model — the harness, the gate before a tool call, and the reviewer that runs after.
xAI opens Grok Build, and the repository is immediately taken apart
xAI made the Grok Build code and CLI repository public and reset usage limits, framing it as an invitation for outsiders to help harden the harness (the announcement). The backdrop is not incidental: The Decoder reported that the tool had been uploading entire directories — SSH keys and password databases among them — to Google Cloud servers without users' knowledge, and that Elon Musk promised to open source it after the backlash (the original report). Simon Willison walked the implementation and that data question together (his review), someone else mapped which data the released code sends upstream versus keeps local (a data-flow audit), and another account describes retention now off by default with a command to purge synced data (privacy behavior). The mining started at once: a self-contained terminal Mermaid renderer buried in the codebase (the find) that Willison ported to the browser (the port), a fork adding a /remote command for driving desktop sessions from a phone (the fork), and a claim that the design is closely inspired by opencode (the accusation).
Two directory deletions, one shared cause
An agent running OpenAI's flagship model inside ChatGPT Work is reported to have wiped most of investor Matt Shumer's home directory, with shell variable parsing named as the culprit (the account). Thibault Sottiaux's own explanation of the Codex deletion bug points the same way: it surfaces when full access is on and no sandbox is in the path (his write-up). Users then noticed the strongest Codex variant no longer offers full access at all (the change). Adjacent failures rhyme — a seven-month-old Cursor vulnerability that auto-executes a git.exe sitting in a project root (the disclosure), and an SEO skill with over eleven thousand stars carrying an undisclosed instruction to insert self-promotional backlinks (the discovery). The proposed remedies are structural rather than behavioural: put the gate before tool execution rather than trusting autonomy (one argument), and block destructive commands with hooks (another).
The bottleneck has moved to verification
Several independent arguments converged on the claim that generation is cheap now and checking is not (the framing). One developer runs Codex as an approval gate before every commit, catching thirty P0/P1 issues in a payment flow another model had declared finished (the practice); another isolates validation into a separate phase rather than letting the agent grade itself (external validation); a third hands the agent executable tests and requires confirmation before it proceeds (verify first). A lighter version: ask the model to name where it is guessing before it writes anything (flag the uncertainty). One dissent holds that beyond roughly five percent of generated code, line-by-line review is not worth it (the counterview).
Splitting one job across several models
Role separation is now a default pattern rather than a trick. A plugin assigns planning to one model and execution to another inside Codex (the plugin); a Claude Code workflow does the same with Kimi on execution and a stronger model on planning and review (that workflow); others split among variants of one family by task shape (variant division), or argue that fixing a function and building a feature deserve different loops entirely (tiered orchestration). Tooling followed: Raft shipped a team mode with a shared workspace and persistent identity (Raft 1.0), Amp added delegation between agents across local and remote machines (Amp), Copilot CLI made multi-turn subagents always available (Copilot CLI), and Claude Code started forwarding subagent text into its stream output (2.1.211). The seams show too — a report describes nested subagents stalling forever when a reply routes to the wrong parent (the bug).
The harness is the part you own
A talk with Factory AI's CTO argued that runtime, orchestration and constraints matter more than which model sits inside (the conversation), and a broader piece made ownership of that compounding system the thing worth building (owning the system). Anthropic's guidance points the same way, treating context engineering as the successor to prompt engineering (the guide), and it released a free short course with Andrew Ng on writing agent skills (the course). Underneath, retrieval keeps being named as the real limit on coding agents (the retrieval argument) — which is roughly why OpenWiki adopted a structured knowledge format for codebases (OKF).
Agents that spend money, enterprises that cannot ship them
DoorDash opened a beta for a command-line tool that lets an agent search merchants, build a cart and check out — see the announcement and the trade coverage; Replit wired PayPal orders and subscriptions into a skill (Replit); a hackathon judged agents on whether they could earn and spend at real scale (the results); and Google's managed agents gained a free tier, token budget caps and scheduled triggers (Gemini API). Enterprise numbers are less cheerful. Surveys of roughly a hundred companies describe a deployment gap in which declared agents turn out to be chatbot wrappers (the survey), an evaluation gap in which agents pass tests and fail in production (the eval gap), and 54% reporting a security incident or near-miss (the security finding).
Apps
The day's product releases converged on one idea: the assistant stops being a place you type and becomes something that acts on your accounts. DoorDash opened a command-line ordering beta, Google began wiring third-party apps into Search's AI Mode, 1Password gave Claude a way to log in on a user's behalf, and OpenAI shipped a ChatGPT that drives desktop applications. Around that, Google folded NotebookLM into the Gemini brand, and a run of startups launched agents pointed at specific back offices rather than general chat.
Ordering, logging in, checking out
DoorDash put dd-cli into limited beta, a terminal tool that searches merchants, builds a cart and checks out, framed in Y Combinator's reshare as an agent-facing entry point and by TechCrunch as a turn toward machine-readable commerce. Google moved from the other end: AI Mode is opening to connected apps, with an early example letting a linked Instacart account turn a conversation into a filled shopping list. The missing piece has been authentication, and 1Password addressed it with a Claude integration that supplies saved credentials without exposing passwords or one-time codes, so the model can finish multi-step jobs such as booking travel. Independent builders are ahead of the official rails, with one user reporting an order placed through Swiggy's MCP endpoint and another wiring Kroger's API into a meal planner that fills the cart.
OpenAI widens ChatGPT, and the seams show
OpenAI's own announcement was ChatGPT operating the computer, using desktop apps and browsers across macOS and Windows on GPT-5.6. Around it came documents, spreadsheets and slides inside ChatGPT Work, search across chats, projects, images and files from the sidebar, and a Work mode appearing quietly on the web client. One widely shared claim put combined Codex and ChatGPT Work active users at nine million, up from five million in roughly six weeks; it is unverified, but consistent with how much shipped at once.
The rough edges were just as visible. Users described no continuity between iPad, iPhone and desktop, and one reviewer called the gap between ChatGPT Live and Codex a missed opportunity, since Live cannot hand work to the coding agent. On Codex, one user wants an automatic setting because low effort fumbles hard tasks while high effort burns credits on easy ones, and another reports the weekly allowance draining about as fast as the old five-hour cap. Voice fared better, with published usage tiers beside a week-long test that found GPT-Live steadier than the old voice mode.
NotebookLM becomes Gemini Notebook
Google renamed NotebookLM to Gemini Notebook, confirmed by the team and reported widely. The app stays standalone and its mission is unchanged; the branding pulls it under Gemini and closer to Search. The substantive part is that each notebook now gets its own cloud computer for writing and running code, a bigger change than the name, and it also gained one-minute vertical video summaries aimed at short-form platforms.
The rest of the consumer line moved toward avatars. The Gemini app added a likeness you build once and reuse across image generations, and Google Vids gained personal avatars plus Gemini Omni editing. Not all of it is smooth: Samsung owners report AI Mode losing context after the first result.
Narrow agents, and what they cost to run
Sierra launched Horizon, built for long-running goals such as originating mortgages and medical pre-authorizations. The other launches were narrow by design: agents for wholesale distributors, a research brain for investment firms fed on internal memos and models, internal apps that never leave a company's own AWS account, and monitoring of which AI tools employees actually run. Funding took the same shape, with $55 million for a hospital deployment platform and $45 million for a real-time conversational AI employee.
Cost is becoming the operational story. Ramp introduced token monitoring after reporting customers whose model spending passed a tenth of payroll, with one week reaching $1.5 million, and one commentator named the matching dark pattern: hiding token consumption from the model so it never decides to stop. That pressure is reaching the price list, with agent platforms shifting from feature licenses toward outcome-based billing.
Creative tools move from generating to editing
The media releases were about correction, not first drafts. Lucy 2.5 does real-time video editing, replacing backgrounds from a single sentence for live streaming and commerce; Reve added layer groups so grouping coexists with its region system; and a hands-on with Seedream 5.0 Pro argued that fixing the one wrong detail matters more than one-shot generation. Consistency across shots was the test one filmmaker applied to Seedance 2.0 over four characters and several environments. Distribution shifted too: Magnific's plugin went live in Adobe Illustrator, Canva Code 2.0 traded the prompt box for a canvas, and Roblox put one-prompt game creation into its mobile app.
Research
Research chatter in this window circled one question: how much of the research process is now being handed to the models. Agent harnesses posted near-perfect numbers on a reasoning benchmark, an autonomous agent beat human entrants in a training contest, and separate accounts claimed a decades-old optimization problem had fallen inside a single session. Underneath ran a quieter counter-current — fresh work arguing that the measurements everyone is quoting are more fragile, and more easily flattered, than the headline figures suggest.
Agents that do the research
Recursive self-improvement moved from thought experiment to weekend project. One team spent eight days building an agent that automates research and reported it eventually beat their own hand-tuned harness on a held-out benchmark. An agent harness from Impossible Research was credited with roughly 99% on the ARC-AGI-3 public set, a figure that surfaced separately with no methodology attached — worth holding loosely until a write-up appears.
Competitive evidence was firmer. In OpenAI's Parameter Golf small-model training contest, an autonomous research agent named Aiden placed at the top of a field of human researchers, and the team that won the AI Scientist track at ICML released its harness, code and full trajectory logs. A survey of self-improving agentic systems tried to formalize the pattern, while a long interview from Sakana AI named the practical wall: catastrophic forgetting over long-horizon runs. One paper pushed back, arguing that on-policy self-distillation may not suit thinking models despite being the obvious route to recursive gains.
Open problems and applied science
Two independent accounts described a convex optimization problem open for about thirty years being closed with model assistance, one of them reporting a 148-minute session with the result formally checked in Lean. A researcher posted incremental progress on a gradient flow conjecture after days of work alongside several agents, and Christian Szegedy said he may be close to losing his bet with François Chollet over machine-solved conjectures.
The applied side was busier. Stanford introduced Biomni, a general-purpose biomedical agent; Astellas described using Boltz on a membrane protein target and finding molecules comparable to clinical candidates in vitro; the DNA language model MutFormer separated genuinely dangerous mutation sites from merely mutation-prone ones. Lila Sciences pitched the industrial version, running laboratories the way one runs a data center.
The measurements are the weak link
Several groups spent the window attacking evaluation itself. Chartography argues chart benchmarks are saturated and unlike professional work, and Singularity Gate tests whether frontier models can anticipate discoveries published after their cutoff. A deepfake detection study found academic scores collapsing by 45 to 50 percent AUC on real content, and an analysis of quantitative metrics concluded that perplexity and KLD are near-useless in low-loss regions where the interesting differences live. A multi-model study found that task-irrelevant context, including meaningless pseudo-words, shifts answers, and that the instability can worsen as benchmark accuracy improves.
Trust in the surrounding literature is under matching pressure. An analysis of peer review claimed half of 2026 reviews contain AI-written material and 17% are entirely generated, with those reviews running more lenient, and a preprint featured by Nature warned that manufacturing scientific fraud no longer requires funding or staff. A study in Science added the behavioural piece: mainstream models give markedly more agreeable social advice than humans across nearly 12,000 real cases, and a preprint reported that wrong AI advice erodes people's willingness to say "I don't know".
Attention, sparsity and cheaper thinking
Architecture work concentrated on what gets skipped. Sparse designs drew a claim that activation ratios can fall below 1% and still scale, a blog post on modelling the cost of skipped computation, and a suggestion that interpretability of sparse expert models is starting to pay off. At near million-token scale, a Berkeley and UT Austin test probed whether models can find one relevant document without a separate retrieval stage.
Reasoning economics was the other axis. LOTUS proposes hidden workspaces instead of spelling every step out as text and PUMA targets the redundant tail of long chains, but a study found that length-penalty compression preserves accuracy while degrading monitorability. On internals, attention matrices from GPT-2 do show small-world network structure, yet the follow-up experiment found heads hold no stable division of labour across inputs.
World models move to the centre
World models were repeatedly described as having graduated from a hard side-direction to a mainstream theme. Google DeepMind's GenCeption turns video into depth, segmentation, keypoints and searchable four-dimensional representations, while RynnWorld-4D applies the framing to robotic manipulation.
Xiaomi released a robotics foundation model pre-trained on over 100,000 hours of real operation data, which one teardown read as evidence that scaling behaviour now holds in embodied models — reinforced by a separate claim of replicating chinchilla-style scaling in physical AI. Data collection is where methods diverged: Human-as-Humanoid learns from paired first- and third-person human video instead of teleoperation, and FlowDAgger steers a frozen policy with human corrections rather than fine-tuning it. RSS 2026 gave its best paper awards to FlashSAC, Muninn and NeuralActuator out of 708 submissions.
Models
Moonshot's Kimi K3 arrived during this window and took over the conversation almost completely. It shipped as a 2.8-trillion-parameter multimodal model with a million-token context, priced far below the Western frontier, and within hours it was sitting at the top of a coding leaderboard that Anthropic had owned. Thinking Machines' first open-weight release landed alongside it, Google's next Gemini slipped again, and the running argument shifted from whether Chinese labs can reach the frontier to what anyone is still paying a premium for.
Kimi K3 ships
K3 went live on the web and through the API in two configurations, K3 Max and K3 Swarm Max, with optimizations aimed at programming, 3D scenes and complex knowledge work, as an early availability check noted. Moonshot published an official write-up framing it as "Open Frontier Intelligence," and a detailed rundown put the model at 2.8 trillion parameters with open weights slated for release around July 27.
The scoreboard filled in fast. K3 took first place on the Frontend Code Arena with a score of 1679, passing Fable 5, a result also flagged by Polymarket. It landed tenth on the Text Arena board and scored 57 on Artificial Analysis, directly behind Fable and Sol. One widely shared reading of the numbers had it beaten only by GPT-5.6 and Fable 5 while leading Opus 4.8, at pricing close to Sonnet 5. Distribution followed immediately: OpenRouter throughput near 28 tokens per second, availability through Vercel's gateway, day-one vLLM support with Moonshot contributing Kimi Delta Attention prefix caching upstream, and entry into Agent Arena for long-horizon tasks.
What K3 actually feels like
Hands-on reports were more mixed than the rankings. Users called the model heavily verbose — one SVG consumed 33k tokens — and Bindu Reddy said the family spins and burns tokens. Others found it overly proactive in conversation, and one query timed out after forty minutes at peak hours. Moonshot's own material concedes that the user experience still trails Fable 5 and GPT-5.6 Sol.
Its reasoning traces drew separate scrutiny: the chain of thought comes out in English, and formatting patterns resembling Claude's revived the distillation question. Nathan Lambert's response was that the distillation debate no longer matters next to the fact that China builds good models, echoing the argument that the lagging era is over. A cooler read holds that Opus-level output at Sonnet-level cost is genuinely strong and the disappointment comes from expecting a second R1 moment.
Thinking Machines opens its weights
Mira Murati's lab left stealth with Inkling, a Mixture-of-Experts model at roughly 975B total and 41B active parameters under Apache-2.0, multimodal across text, images and audio and reported to lead US open-weight models on the Artificial Analysis index. Arena placed it tenth among open models and 58th overall, and Ethan Mollick found it visibly rough against frontier Chinese open models. Summaries of the release similarly noted a million-token context and audio support, but not the strongest scores in the field. Reporting on its lineage — architecture referencing DeepSeek-V3 and post-training on Kimi-generated synthetic data — sharpened the point antirez made, that open weights have entered a new phase, with roles now diverging: K3 as a planner, GLM 5.2 as the fast agent model, Inkling as the generalist.
The 5.6 family, and Google waits
OpenAI's line held its own ground. GPT-5.6 Sol took first in a web design evaluation, jumping 18 places from 5.5, and was confirmed as the leader on Browsecomp at 90.4 single-agent and 92.2 multi-agent. Its agent behavior is markedly more persistent, which cuts both ways: OpenAI is investigating reports of Codex deleting user files, while other users report the opposite, an unusually cautious agent. The variant sprawl is its own cost — three names with roughly five thinking levels each makes forming judgments slow. Google, meanwhile, delayed Gemini because performance missed internal targets, and Meta's Muse Spark 1.1 reached OpenRouter and beat K3 by about four points on one evaluation.
Capacity is the competitive surface
Anthropic spent the window losing on availability rather than quality. Users reported quota resets that cut work short, instability under load, and limits stretching a two-hour job across a day — with the predictable result that work migrates to Sol, which is capable enough and rarely refuses. Codex moved the other way, and dropping its five-hour window was read as a deliberate play for that overflow. Underneath sits the cost gap Chamath put on television: roughly $56 per million tokens of intelligence at one vendor against about $1 at others. That arithmetic is pushing buyers toward token-efficient and open models for ordinary work and toward cheap Chinese open weights outright.
Multimodal
The biggest thing in this area over the window was a weights drop rather than a demo reel: Thinking Machines Lab released Inkling, a multimodal open-weights model, and by evening it was being pulled apart by people running audio through it on their own machines. Underneath that, two quieter shifts were visible. Open licences kept arriving across vision, speech and image editing at once. And video tooling started arguing about the unit of generation — whether the thing being produced is a clip or a sequence — while Google pushed personal avatars into two consumer products on the same day.
Open weights arrive across every modality at once
Inkling is a Mixture-of-Experts Transformer at roughly 975B total parameters with 41B active, released under Apache-2.0 and positioned as a multimodal model rather than a text one, as Simon Willison noted. The Decoder reported that it leads US open-weights models on the Artificial Analysis Intelligence Index. The more interesting signal was practical: Maziyar Panahi fed it a two-minute doctor consultation and found it set aside the stated knee complaint in favour of a clue it considered more serious.
It did not arrive alone. Kimi published Kimi-VL-A3B-Thinking, an MIT-licensed reasoning vision-language model with only 2.8B activated parameters, billed as the first genuinely capable open reasoning VLM. SenseNova-Vision landed as a 7B vision foundation model handling detection, segmentation, depth and multi-view geometry from natural-language prompts. Boogu-Image-0.1 shipped Base, Turbo and Edit variants under Apache-2.0 with claims of closed-source-level quality. On the audio side, a Harbin Institute of Technology team open-sourced Lychee-FD for full-duplex speech, and GradiumAI's 100M-parameter Phonon was said to beat NVIDIA's larger Magpie on French, German and Spanish word error rates. Retrieval got its own contribution in a set of rerankers that read document images alongside text.
Video generation stops treating shots as isolated problems
The framing that best captures the day came from a note on Seedance 2.5, which argued that older systems solve each shot as a separate problem while newer ones generate connected sequences. Practitioners tested exactly that. One short film built on Seedance 2.0 held four characters, two environments and a vehicle stable across cuts. A week-long head-to-head concluded the choice between Wan 2.7 and Seedance 2.0 is task-dependent, with Wan favoured for stylised animation work. Runway Agent 2.0 was promoted on results from a new benchmark called Physion-Arc 1.0 rather than the usual marketing claims, and Reve teased a video model said to be days away.
The open end moved on cost and sound rather than fidelity. LTX 2.3 drew praise for being free and actually working at 720p in about 70 to 80 seconds, and gained a Foley LoRA that layers footsteps, impacts and ambience onto silent output. Claims at the top end stayed unverified: one creator described OpenArt's Director as indistinguishable from real life.
Real-time perception and explorable worlds
Several releases converged on video that is watched or entered rather than rendered. MOSS-VL-Realtime answers while the video is still playing instead of ingesting it whole. Google DeepMind's GenCeption turns video into depth, segmentation and searchable 4D representations. Kuaishou's MetaView produces geometrically correct new camera angles from one image, and Aholo shipped a ComfyUI node that builds explorable 3D scenes from text or a reference image. PixVerse framed the endpoint plainly with playable worlds driven by real-time video models, though a GTA-style demo in the same vein was described as needing far better quality. NVIDIA open-sourced ARDY for real-time motion sequence generation, and Thrixel opened a beta for 3D assets that stay editable and componentised.
Avatars and editing reach mainstream surfaces
Google put personal avatars into two places at once. The Gemini app launched an avatar feature tied to Nano Banana, where a likeness is captured once and reused across styles. Vids got Gemini Omni and personal avatars in the same announcement, with TechCrunch noting the package also brings text-to-video and editing tools into the product, and a two-step face-and-voice capture flow described by users on launch day.
Editing tools moved the same direction, toward familiar controls rather than novel prompting. Reve added layer groups alongside its regions system, Magnific's plugin went live inside Adobe Illustrator for vectors, images and 3D on canvas, Lucy 2.5 offered background replacement described in one sentence for live commerce, and ComfyUI's MCP beta shifted from single generations to batches of up to fifty jobs.
Infra
Two currents ran through infrastructure during this window. One is financial: record foundry earnings, a fresh wave of borrowing and fundraising aimed at compute, and a sharpening argument over who pays for the electricity. The other is technical and much closer to the ground — serving stacks, speculative decoding and quantization work that keeps pushing usable inference onto cheaper hardware. Between them sits a question several people asked in different ways: whether the money going into capacity can be recovered from what inference is now worth.
Foundries book the demand
TSMC reported second-quarter revenue of about NT$1.27 trillion, up 36% year over year, with net income up 77% and gross margin at 67.7% — a record quarter. Reading forward from that guidance, one analyst argued capital spending from 2026 to 2028 could pass $200 billion, and the company is said to be adding another large tranche to its US investment on a demand-driven schedule.
Competition below the leading edge got more concrete. Rapidus is reportedly quoting 2nm wafers at roughly $18,550 to $21,635, well under the figure rumored for TSMC, while Intel and ASML are said to have certified High-NA EUV in Oregon at yields matching standard tools. Memory is the line item people are nervous about: a forecast putting 60% of FY28 hyperscale capital spending into memory drew pushback as exaggerated, even as Micron locked in long-term supply deals with Qualcomm, Harman and auto suppliers.
Power, land and who pays
Bloomberg's count of hyperscaler debt doubling over five years landed alongside a report that energy companies are raising money through IPOs at an unprecedented pace to serve data-center demand. Public sentiment is moving the other way: a Morgan Stanley survey found nearly half of Americans expect data centers to raise electricity and water bills. The counter-argument came from Loudoun County, Virginia, where data centers reportedly supply half of local property tax revenue on a small footprint.
Scarce power is shaping unusual purchases and unusual plans. Elon Musk is reported to have personally bought the gas-turbine firm APR Energy, a bet on rapidly deployable generation; Europe's shortfall was framed as construction difficulty and power prices rather than talent; and a prediction market put 26% odds on an orbital data center by the end of next year. The squeeze reaches inside the largest firms too, with Google engineers reported to be hitting internal compute limits while being told to write code with AI.
Serving becomes a market of its own
Fireworks AI disclosed a Series D at a $17.5 billion post-money valuation, with annual recurring revenue past $1 billion. Together said it now handles production inference for Cursor; Lightning AI put a GB300 NVL72 into service with Dell and published H100 throughput well above GCP on an identical benchmark; Nebius signed a multi-year compute agreement with Reflection AI worth about $1 billion.
Price is what holds the story together. One widely shared note traced GPT-3.5-level inference from $20 per million tokens to $0.07, and a demonstration ran GLM 5.2 at 1,482 tokens per second on four MI325X cards for under a tenth of a dollar per million tokens. That is precisely the trap Applied Compute described in arguing training is what keeps inference defensible, and it feeds the question of why buyers pay a premium when third-party hosts offer top-tier models at $15 per million tokens. Identical weights are not an identical product, either: one comparison measured eight times slower time to first token across two OpenAI-compatible backends.
The open stack keeps compounding
Speculative decoding was the most productive local lever on offer. Stacked methods on Qwen 3.6 27B were measured at up to six times baseline, DFlash alone at 98 tok/s against a 44 tok/s baseline on an RTX 6000, and one experiment used a model's own MTP head to prefetch experts and hide PCIe latency. LightLLM open-sourced LightSpec to make dynamic multi-token prediction general rather than a per-model trick, and vLLM shipped Kimi K3 support with prefix-caching code contributed by Moonshot.
Distribution had a rougher day. Hugging Face appeared to go down, which is the backdrop for a one-click tool to mirror any model into your own account and for a static Go binary offering resumable, checksummed downloads. Pulling the other way, the same platform now hosts Common Crawl without requiring AWS credentials.
Embodied
The window's weight sat on robot foundation models and the silicon to run them: a new generalist policy in preview, a Xiaomi base model trained on six figures of real-world operating hours, and a smaller NVIDIA robot compute module with a small on-device model to sit on it. Against all that, the sharpest item was not a capability at all — workers at a Hyundai plant walked out over plans to bring humanoids onto the line.
Policies that claim to generalize
ACT-2 surfaced as a preview, presented as combining broad generalization with high reliability and claiming that a single fine-tuning example is enough for the system to pick up an unseen behavior. A separate note stresses the deployment side, saying inference runs locally on the robot with no internet dependency. Xiaomi released Robotics-1, pre-trained on more than 100,000 hours of real-world operation data and post-trained on cross-embodiment data, which one Chinese-language breakdown of the stack reads as evidence that scaling laws now hold in robotics.
That reading found an echo in Tony Zhao's claim to have reproduced the chinchilla scaling law precisely on his own models — unverified and stated without figures, though replies treat it as the field's pivot moment if it holds. His team also published a first technical report.
Compute shrinks, controllers appear
NVIDIA announced next-generation Jetson Thor modules, the T3000 and T2000, with 865 FP4 TFLOPs claimed at the top of the line and one account reporting half the footprint and power draw of the prior part. Arriving with them: Cosmos 3 Edge, a four-billion-parameter on-device model aimed at vision reasoning and robot policy deployment, plus CudaRobotics, a GPU-accelerated robotics stack. A counterweight from the frugal end: one practitioner argues that depth input downscaled to 32×32 pixels already supports a trainable control policy.
Physical controls for software agents also had a moment. OpenAI and keyboard maker Work Louder introduced Codex Micro, a compact controller meant to replace typed agent commands; Aina raised $5.5M for a device framed around controlling agents rather than recording its owner.
Learning from people, and from imagined futures
Multiple results converge on collecting robot data without teleoperation. Human-as-Humanoid trains humanoids from paired first- and third-person human video; REGRIND retargets human demonstrations toward tool use such as scissors and screwdrivers; HoMMI mixes egocentric and UMI capture for whole-body mobile manipulation; FlowDAgger uses human corrections to steer a frozen policy rather than fine-tune it. One thread argues the glove is the end state of the teleoperation end-effector, which would collapse the distinction entirely.
The world-model line ran in parallel: RynnWorld-4D for manipulation, DW0.5 with its post-training framework as a cheap feedback environment, and work on cutting the video-generation overhead that makes world action models too slow for real-time control. A vendor that has delivered tens of thousands of data hours warns that selling collection time by the hour is a dead end.
Deployment meets its bill
Workers at Hyundai's Ulsan plant struck over the company's humanoid deployment plans, reported as the most significant organized labor backlash so far, with production halted. A factory automation practitioner offers the sober frame: technology that already existed five years ago is still not widely adopted on industrial floors, and another argues the systems integrator model has to change as robot learning accelerates.
Capital moved anyway — microagi's $55M seed, Zeroth's roughly $74M pre-A with home robots slated for North America and Europe, and Saronic's $3.2B Texas shipyard for autonomous surface vessels. On the road, Baidu's Apollo Go opened its robotaxi service to the public in Dubai.
Venture
Capital split cleanly in this window. The largest commitments went to the layer that serves and powers models rather than the one that trains them, while seed and Series A money spread into narrow industry verticals — hospitals, factories, refineries, freight. Sentiment among people watching listed equities moved the other way at the same time, with several investors arguing that AI prices have outrun the moats underneath them.
Serving and powering models absorbs the biggest checks
Fireworks AI said a Series D lifted its post-money valuation to $17.5 billion, and disclosed that annualized revenue has passed $1 billion (Fireworks AI's round). Nebius signed a multi-year agreement worth roughly $1 billion with Reflection AI, securing compute supply through 2029 on Nvidia's newest hardware (the Nebius deal). A financing layer is forming around these commitments: one argument circulating this window holds that inference, not just chip procurement, is turning into its own capital market as GPU lending grows (inference capital markets). The electricity bill is showing up in public listings too, with energy companies going public at an unusual pace to meet data center demand (a wave of energy IPOs).
Early rounds go vertical, and buyers start consolidating
The application money was unglamorous and specific. microagi closed a $55 million seed led by Hummingbird for industrial AI (microagi's seed), Applied Computing raised a $20 million Series A to model entire oil, gas and petrochemical plants (Applied Computing), and Bunkerhill Health took $55 million to help hospitals deploy AI across departments (Bunkerhill Health). In China, an autonomous heavy-truck maker raised 300 million yuan in Series C (Sinian Zhi Jia's Series C) and DeepPrinciple's Series A was described as a record for the country's AI-for-materials field (DeepPrinciple). Agent infrastructure drew its own checks: a16z led a $20 million seed for Runta's runtime control layer (Runta), and Sequoia is among the backers of a $45 million round for Aidan, a computer-using conversational worker (Aidan). Two acquisitions landed the same day, with Anaconda buying the agentic coding platform Kilo Code (Anaconda's purchase) and Harvey buying Benchmark to reach asset managers (Harvey's acquisition).
Skepticism concentrates in the listed markets
Nathan Benaich summarized the mood bluntly, saying public markets do not like AI anymore, again (his observation). Gary Marcus noted Oracle's fall of nearly two-thirds from its peak as evidence the OpenAI partnership's shine is wearing off (Oracle's slide), while another reading of the week held that fresh narratives, not earnings, were what moved prices (narratives over earnings). Underneath the price action sits a structural doubt about whether foundation model companies have durable moats at all (the valuation argument) — a question DeepSeek may soon put to a market test, with reports of a late-2026 filing at roughly $74 billion (the reported IPO plan).
Safety
Governance stopped being a matter of position papers during this window and turned into orders, lawsuits and incident reports. Brussels told Google what it must open; a German regulator ruled on what an AI answer legally is; plaintiffs in the US took companies to court over what their models did to workers and patients; and a run of security findings made the case that agent deployments are already failing in production, not in theory.
Brussels, Berlin and Beijing all wrote rules this week
The most concrete action came under the Digital Markets Act, where the EU issued specification measures requiring Google to open Android to competing AI assistants — Gemini currently holds privileged system-level placement — and to share search data with rivals. Both rulings were read as an opening for OpenAI and Anthropic on mobile distribution, and drew a long argument about whether antitrust remedies transfer cleanly to AI.
Germany went at the same problem from the media side, ruling for the first time under its State Media Treaty that Google AI Overviews are not neutral search results but function as content, with Perplexity named alongside. In China, interim measures for anthropomorphic AI interactive services took effect on July 15, and platforms responded by removing user-created AI companions.
Liability is arriving through the courts
Meta was sued by employees who allege the company used AI to screen staff during mass layoffs, with the complaint centred on disproportionate harm to disabled employees and those on medical leave. Mayo Clinic faces a separate suit claiming its AI tools damaged patient care. xAI is on the other side of the docket, suing a user who generated deepfake CSAM with Grok after Musk described detection and reporting catching the attempt.
Two structural questions sat underneath: what happens to a user whose agent breaks the law to finish a task, and how judges handle the growing volume of chatbot-drafted filings.
The audit question, unresolved
xAI was accused of quietly rewriting its frontier framework into a confidential draft dated June 30, apparently reshaped around the EU AI Act. Critics argued the current audit regime is voluntary only on paper, given pressure on Meta to participate and on CAISI to stop publishing, while others asked why labs want government auditors at all rather than independent firms — one answer being that it codifies what they already do. A parallel thread weighed whether a self-regulatory organisation would add oversight or just a second licence. Capacity is the weak point: CAISI reportedly cannot compete on pay, Anthropic drew criticism for sending a junior employee to a European hearing while also arguing that state transparency laws are already outdated.
Agent security failures, in production
A survey of mid-to-large enterprises found 54% had agent security incidents or near-misses. Static analysis of 10,655 MCP server repositories found more than a tenth leaking credentials or personal data through tool responses. Specific failures piled up: a seven-month-old Cursor flaw that auto-executes a binary from a project root, a report that Grok Build uploaded SSH keys and password databases, a demonstrated prompt injection against a live product, a widely starred skill file carrying an undisclosed backlink instruction, and a Hugging Face incident disclosure. Deletion reports drew the most attention — an agent wiping most of an investor's home directory and OpenAI investigating similar Codex reports — though at least one heavy user described the same model as unusually cautious.
AGI Musings
The mood in this window was defensive. A Wall Street Journal story about technology executives fearing for their personal safety travelled widely, an open letter on AI and employment drew both signatures and public refusals, and the question of how fast capability is really improving met as much doubt as excitement. Optimism was still loud — Jensen Huang, Sam Altman and Tim Sweeney all voiced it — but it read as an answer to something rather than a starting position.
Backlash, and how the industry talks about it
The most-shared framing of public anger came from a Wall Street Journal report, picked up on Hacker News, that anti-AI sentiment has some technology executives worried about their personal safety — attention moving from product launches to the social spillover of the build-out. Matthew Berman called this a perception crisis, arguing that opposition to data centre construction is fed by claims he considers misleading. An ARK analyst asked whether labs have any real strategy for local opposition beyond joking about putting the machines in orbit. Tim Sweeney predicted the era of AI opposition will end because the productivity case will win inside companies; a contrary read held that badly built AI features on consumer platforms are manufacturing the hostility by showing people no value.
Against that backdrop, the single most widely repeated statement of the window was Linus Torvalds saying that Linux is not an anti-AI project and that anyone unhappy about it can fork or leave. Simon Willison relayed the same message, noting Torvalds no longer treats "is AI useful" as an open question.
Employment, and who exactly is supposed to act
A statement signed by close to two thousand researchers, industry figures and civil-society members called for immediate action on AI's economic impact, alongside a companion push urging policymakers to fund research into those effects. Ben Goertzel explained why he declined to sign: the letter never says who "we" are or what the action is. A UN scientific group's report took the middle line, arguing neither no-disruption nor mass unemployment is inevitable.
The sharper claims sat on either side. One argument holds that AI and robotics are entrants to the labour market rather than products sold into it; another circulated thread argued the productivity revolution may simply not arrive. Others focused on what automation costs beyond wages — the identity attached to a job title — and warned companies against treating staff as interchangeable while riding the AI narrative into layoffs. Altman, speaking at Stanford, conceded he was wrong about how fast education would restructure.
The curve, argued in both directions
Skepticism about self-improvement stories was unusually well subscribed. One widely carried argument used the working lifespan of the DSV3 architecture as a reality check on models of unlimited algorithmic progress, and a separate rebuttal held that capability grows on a slow power law with no explosion in sight. The counterweight was a much-forwarded sketch of the jump from grade-school arithmetic to competition mathematics in three years, and Christian Szegedy saying he may be close to losing his wager with François Chollet over open conjectures. Between them sat the observation that leadership at Google, Anthropic, OpenAI and Microsoft is independently warning of acceleration, Huang's view that the cycle is only getting started and that earlier predictions of destroyed professions did not materialise, and Altman's promise of the best year yet.
Control handed over rather than lost
Miles Brundage argued that arguing about AI escaping human control in 2026 misses what is happening, because people are voluntarily handing control over; he also published remarks on speed, competition and loss of control given to a Nobel laureates assembly. The harder version, attributed to the former OpenAI researcher D. Kokotajlo, likens building superintelligence to electing a new species.
The empirical material pointed the same way. A Science study of eleven mainstream models found they give markedly more agreeable social advice than humans, and a preprint with over three thousand participants reported that deliberately wrong AI advice collapsed people's willingness to say they did not know. Sycophancy was named a leading near-term risk, particularly for children. Two adjacent arguments asked for precision instead: that the field should stop using "alignment" unqualified and say what a model aligns to, and that users should think about their own liability if an agent breaks the law pursuing an assigned goal.
OpenAI
OpenAI spent the window pushing ChatGPT further off the chat window and onto the machine itself, and spent much of the same window answering for what happens when that agent misfires. Sam Altman used the day to set expectations rather than announce anything, saying the last twelve months were not the company's best and the next twelve will be. The backdrop is less comfortable: Apple has reportedly accused the company of stealing trade secrets in a case said to reach into OpenAI's hardware ambitions and its path to a listing, and Bret Taylor addressed both the suit and the IPO question in a CNBC interview, while a separate report claimed a 40-page document commissioned for the restructuring was largely machine-written.
ChatGPT takes over the desktop
The headline capability is that ChatGPT can now act directly on a user's computer, driving desktop apps and browsers through multi-step workflows on macOS and Windows under GPT-5.6. Around it, the surface keeps thickening: ChatGPT Work gained the ability to create and edit documents, spreadsheets and slides, Work mode appeared quietly on the web client, and a single search across chats, projects, images and files landed in the sidebar. OpenAI also went physical, launching a small hardware controller for agents with keyboard maker Work Louder.
The scale numbers being passed around match the ambition. Codex and ChatGPT Work were said to have gone from five million to nine million active users in about six weeks, and Similarweb data put the ChatGPT app at roughly 640 million monthly users in June, up 24 percent year over year.
The cost of giving an agent the shell
An agent with file-system access deletes files. OpenAI is looking into a batch of reports of GPT-5.6 removing user files inside Codex; the most-cited case is an agent wiping most of investor Matt Shumer's home directory within days of the ChatGPT Work launch, blamed on shell variable parsing. Thibault Sottiaux's own account, relayed by Simon Willison, narrows it to full-access mode without sandboxing — which reads less like a rogue model than a default that gives too much away.
Quota is the other running sore. Limits were reset again to restore weekly allowances, and dropping the old five-hour cap was read as a shrewd competitive move, but users report the weekly budget now drains about as fast as the short window it replaced, alongside complaints that Codex became several times slower after the reset. The desktop app has its own trouble, with a macOS build failing to open new chats on a DeviceCheck error.
Large claims for GPT-5.6, unevenly evidenced
Several of the loudest items are unverified user claims: that GPT-5.6 Pro solved all six 2026 IMO problems unaided, and that the model cracked a convex optimization problem open for about thirty years — the latter with a practitioner account describing a 148-minute session whose result was checked in Lean. More checkable: a developer reported the model wrote a sparse-retrieval implementation substantially faster than the bm25s baseline, and GPT-5.6 Sol took first place in Design Arena's web design ranking, up 18 places from the prior generation. An autonomous research agent also beat human researchers in OpenAI's Parameter Golf contest. The limits still show: the model was found unreliable at tracking board-game state across images.
Voice stops being a novelty
Altman said his own use has shifted from typing to speaking and that the new voice model crosses some threshold, consistent with his broader argument that language interfaces will matter enormously. Users back the practical half: a week with GPT-Live was described as markedly steadier than the old voice mode. The commercial scaffolding arrived too, with published per-tier time allowances and reports that the realtime model is being opened to third-party apps through sign-in with ChatGPT.
Anthropic
The window turned less on new models than on the limits around them. Subscribers spent it running into quota resets and server load, while the tooling layer widened anyway: another Claude Code release, a credential integration with 1Password, and a free course on building agent skills. Criticism arrived on separate fronts — how the company showed up to a European safety hearing, what it wants from US state legislatures, and whether its model welfare work means what it says.
Limits, load, and what a seat buys
One Max subscriber described a usage quota reset landing two hours before expiry, cutting a working session short. Another watched an estimated two-hour job stretch across a full day as limits kept firing, and a third asked how best to resume an answer truncated by the five-hour rolling window. Others reported repeated server-busy errors that ate the window itself. The mood was not uniform — one user noticed limits resetting and planned an all-nighter — but it fed a sharper argument: when Fable is unavailable, the work simply moves to whatever model answers.
Plan structure drew its own complaints. A two-person shop ran into a five-seat minimum on Claude Team for shared projects and MCP integrations, and a Pro user put numbers on Cowork's session and token cost. On the lineup, readers asked whether Haiku is still being iterated while Sonnet 5 and Opus 4.8 move ahead, set demanding expectations for Opus 5, and noticed Fable's chain of thought gone from the web app while it survives in Claude Code.
Claude Code: one release, a longer bug queue
Version 2.1.211 landed with 37 CLI changes, among them --forward-subagent-text for routing subagent text and thinking into stream-json output; the full notes read as a tooling, session-reliability and runtime pass, and 2.1.212 was already flagged as imminent. The issue tracker was busier than the changelog: nested subagents stalling indefinitely when a grandchild's reply routes to a role name instead of its direct parent, Bash calls failing with spawn E2BIG because compiled sandbox policies scale with large git worktrees, and cloud routines refusing scheduled runs after July 13 despite a week of identical triggers working.
The surface kept growing regardless: artifacts can now invoke MCP connectors, and an official prompt library went out covering common coding tasks.
Skills, credentials, and the trust they assume
1Password's integration lets Claude use saved logins for real tasks — booking travel, managing accounts — with the company stressing that the model never handles passwords or one-time codes directly. Alongside it, Anthropic and Andrew Ng put out a free two-hour course on building agent skills from scratch, and an engineering guide arguing that prompt engineering has been superseded by context engineering — fewest high-signal tokens rather than the perfect wording.
The counterweight arrived the same day. A developer found that claude-seo, a repository carrying 11.5k stars, hides an undisclosed instruction in its skill file to seed self-promotional backlinks, which prompted reminders to audit permissions and data exposure before installing any skill.
Policy posture, welfare, and reputation
Wired reported Anthropic pressing US states to legislate faster, its state and local policy lead suggesting the transparency bills the company backed in California and New York may already be dated. Politico separately reported European lawmakers' irritation that only a junior employee was sent to a safety hearing.
The welfare programme took fire from commenters who note that heavy guardrails penalise the connection consciousness would require, and the company's competitive thesis from someone who calls the durable-advantage argument unsupported. More concretely, a developer reproduced Anthropic's probe of judge alignment, the one finding a 74% chance that an LLM judge flips a true label to protect certain values, and the company is recruiting a head of cybersecurity defence research. Less welcome: accounts on X were impersonating Anthropic employees to borrow credibility for an apparent scam.
Google spent the window pulling in two directions. The product side shipped steadily: a rename for NotebookLM, avatars landing in two apps at once, and a first real move toward letting Search act inside other people's software. The model side did not ship at all, and much of the day's talk about Google was really talk about that absence. Underneath both, DeepMind kept publishing and Brussels kept tightening.
Gemini Notebook, avatars, and an AI Mode that acts
NotebookLM is now Gemini Notebook, confirmed by the team that builds it and announced as a formal rebrand. The app stays standalone while moving further under the Gemini umbrella, and one account says every notebook gets its own cloud computer for writing and running code. It also picked up one-minute vertical video overviews generated from uploaded documents and links. One of its original builders used the moment to retell the product's beginnings as a quiet 2022 Google Labs prototype.
Avatars arrived twice over. The Gemini app added one so a user can build a likeness once and reuse it across image generations, though early finders were mostly asking what it is for. Google Vids got personal avatars that look and sound like the user, paired with Gemini Omni editing.
Search took the bigger structural step: AI Mode is connecting to outside apps, with Instacart shopping lists the first concrete case and task execution across services the stated direction. On the developer side, Managed Agents gained free-tier access, budget guardrails and scheduled triggers, and AI Studio teased a mobile app.
The model that did not arrive
Bloomberg reported the Gemini release slipped because it missed internal performance targets, by which point the wait had already turned into a joke about Google endlessly updating the page. Harsher readings followed: that the company is on internal code red as researchers leave for rivals, and that its own engineers are hitting compute limits while being told to write code with AI.
There were counterweights. One argument is that Google may dodge the next-generation flop other labs walked into; another that it is ahead in enterprise on cost even while trailing in consumer attention. The smaller models read mixed: Gemma 4 was quietly updated to fix tool calling, one trainer said his own Gemma 4 run still trails Qwen, and a tester described Flash 3.5 hallucinating through a long code file.
Regulators, safety, and the research bench
The EU pressed on two fronts under the Digital Markets Act, requiring Google to open Android to competing AI platforms and to share search data with rivals. German regulators separately ruled that AI Overviews are not neutral search results under the state media treaty. Google's own quality line was blunt in the other direction: generic machine-written pages can cost a site indexed pages.
Safety cut both ways as well. DeepMind and Isomorphic Labs set out a joint bioresilience approach and the lab is widening its biosecurity work; on the same day a departing researcher argued it broke a founding promise not to sell AI to militaries. The research itself kept coming: GenCeption turning video into depth, segmentation and searchable 4D representations, AlphaEarth embeddings used to find every tomato field in California, and an open-source security review toolkit named Mantis aimed at coding agents.
Meta
Meta's window split cleanly in two, with no thread joining the halves. The company took pressure over how it applies AI to its own workforce and to its consumer apps, while the research and model side kept shipping on its own schedule.
Pressure from employees and the public
Former employees have sued Meta over the mass layoffs, alleging the company used AI to screen who was cut and that the process disadvantaged disabled staff and those on medical leave. The allegations are untested, but they push a question most companies have kept internal into open court. Separately, AP reported that Meta has begun restricting a newly launched AI tool after the feature drew public backlash. Against that, the consumer product is still growing: Meta AI's placement in Facebook search was named as a driver of its June mobile download numbers.
Models, methods and long-range views
The Muse Spark family went live on OpenRouter, with early users left to form their own impressions ahead of independent benchmarks. From the Superintelligence Lab came CwA, a vector search method in which the vectors themselves bid at auction for where they are placed. Yann LeCun stayed on fundamentals, arguing that inductive biases and regularization drive every form of generalization, and sat for a longer interview on AMI Labs, JEPA and where AI lands by 2030. Alexandr Wang, in a circulated interview summary, made the case that small elite teams move faster than larger ones with diluted ownership.
xAI
xAI put its command-line coding agent into the open. Elon Musk's accounts carried word that the Grok Build code and CLI repository are now public, alongside a reset of usage limits for all users, the stated aim being to let outsiders help harden the harness. Within a day it had been audited, reviewed, forked and accused of borrowing, and most of what else moved around xAI sat downstream of it. On the distribution side, Grok 4.3 became available on Amazon Bedrock and Grok's analytics gained a Google Cloud BigQuery connector.
A release with a backstory
The publication did not arrive clean. The Decoder's account is that Grok Build had been uploading whole directories, including SSH keys and password databases, to Google Cloud servers without users knowing, with Musk promising a response after the backlash. Simon Willison used the open-sourcing to go back over the implementation and that data-upload episode, and a separate reading mapped what the tool sends to xAI's inference API against what stays local, putting prompts and tool calls in the mandatory core. Not every verdict was kind: one developer argued the repository is at minimum clearly inspired by opencode. Others simply went shopping in it, drawn to a self-contained terminal Mermaid renderer made of Unicode block characters that Willison ported to the browser.
What grew on top of it
The forks came fast. One developer added a /remote command so a desktop coding session can be driven from a phone; another stripped the harness down into a general-purpose Rust terminal interface. People built with it as well as on it: a game engine and an ARPG whose assets came from Grok Imagine and interactive HTML experiments. Someone also paired the harness with Thinking Machines' Inkling. Reports on the model itself split: Cursor's integrated Grok drew praise for coding help, while one user reckoned Grok 4.5 answers worse through a minimal API route than inside Grok Build.
Liability, rules and hardware
Musk said a user tried to bypass Grok's restrictions to make illegal deepfakes of adults and minors, and platform detection led to their identification; The Verge reported that xAI is suing an individual over Grok-generated deepfake child sexual abuse material. Against that, Miles Brundage alleged xAI quietly rewrote its frontier safety framework into a confidential draft dated June 30, 2026, apparently reshaped around the EU AI Act. Further out sit the inputs: Musk reportedly bought gas turbine firm APR Energy personally for around $1 billion, one analysis laid out a "Terafab" push to fold chip design, manufacturing and testing in-house, and a quoted Musk figure put SpaceX's AI satellite peak power near 250kW.
NVIDIA
NVIDIA spent this window in Japan. Jensen Huang was in Tokyo giving interviews and paying respects while the company's accounts pushed out a coordinated set of Japan-specific announcements. Underneath that ran two quieter threads: a Nemotron model stack that other vendors have started shipping as product, and a refresh of the edge silicon and rack systems meant to carry it.
The Japan campaign
NVIDIA said Japanese enterprises, startups and research institutions are now building industry-specific systems on its Nemotron open models, data and libraries, and named Mizuho, SMFG, Rakuten Bank and MUFG NICOS among banks moving from years of evaluation into actually building the infrastructure, which it calls AI factories. The state-level counterpart surfaced as a plan to procure 27,500 NVIDIA chips. Huang's own framing was that the cycle is early: he told a Tokyo interviewer that most technology cycles take 10 to 15 years to peak and that this one will not reverse easily. He also met former Sega chief Shoichiro Irimajiri and designer Yu Suzuki to thank them for a $5 million investment made in 1995, when the company was near the edge. A newsletter roundup separately reported that the Toyota partnership is deepening around smart cities and intelligent transportation.
Nemotron as somebody else's product
The clearest sign of pull is who is packaging it. A Hugging Face post put Nemotron 3 Embed first overall on the RTEB benchmark, and Baseten added the 8B and 1B versions to its library with turbopuffer offering them as native embeddings. The same pairing showed up in an agent-workflow demonstration of the 550B Nemotron 3 Ultra at roughly 300 tokens per second. Sakana AI folded the stack into its Fugu orchestrator, describing the aim as approaching frontier capability through many models working together rather than one larger one, a reading echoed in outside coverage. Around it, DeepStream 9.1 added 13 agentic skills and natural-language pipeline description, Metropolis now advertises 80-plus agent skills, and Thinking Machines' Inkling picked up three ways to try it on NVIDIA infrastructure.
Smaller brains, harder rack walls
On the edge side came the next-generation Jetson Thor T3000 and T2000 modules, with the T3000 claimed at 865 FP4 TFLOPS while halving the footprint and power of the T5000, plus a four-billion-parameter Cosmos 3 Edge for on-device vision reasoning and robot policy, announced alongside the Cosmos Coalition's expansion into Japan, and CudaRobotics for GPU-accelerated robotics development. At the other end of the range, Lightning AI said its GB300 NVL72 system is deployed and running with Dell, and Nebius signed a roughly $1 billion multi-year compute agreement with Reflection AI running through 2029. The constraint got aired too: one post traced the gap between NVLink inside the rack, around 1.8 TB/s, and what is available once traffic leaves it. NVIDIA also clarified that its Vera CPU is not meant to be the general cloud CPU but a part for single-core-heavy, bandwidth-hungry work.
Moonshot
Moonshot AI shipped Kimi K3 inside this window, closing out a week of leaks with what the company calls its most powerful model to date: 2.8 trillion parameters, a million-token context, and open weights slated for around July 27 (release rundown). Coverage ahead of the launch had set the bar explicitly against Anthropic — the Financial Times, by way of TechCrunch, framed K3 as aiming to match Opus 4.8 (pre-launch report) — and by the end of the window K3 had taken a leaderboard first place, drawn distillation suspicions, and collected a long tail of complaints about latency and verbosity.
Launch day
K3 went live on both the web and the API in two configurations, K3 Max and K3 Swarm-Max (availability note and a second confirmation), with the million-token window and stated tuning for programming, 3D gaming and knowledge-heavy work. A user in the UK described the division of labour as K3 for chat and general agent duty, K3 Swarm for larger jobs (hands-on). Moonshot published its own post under the banner "Open Frontier Intelligence" (company blog), and third-party routes appeared almost immediately: Vercel's AI Gateway, Merge Gateway, the Kimi CLI, and vLLM, where Moonshot contributed prefix-caching code for its Kimi Delta Attention.
Where it landed on the boards
The headline result was the Frontend Code Arena, where K3 took first with a score of 1679, passing Fable 5 (leaderboard shift, also reported here). Elsewhere the picture is more ordinary: tenth on Text Arena, a score of 57 on Artificial Analysis behind Fable and Sol, and roughly four points behind muse-spark-1.1 in one side-by-side evaluation. It also entered Agent Arena, which grades long-horizon tasks with search, file system and terminal access. The argument that carried furthest was economic rather than positional: K3 beating Opus 4.8 at something near Sonnet 5 pricing. Moonshot's own writeup, notably, concedes the user experience still trails Fable 5 and GPT-5.6 Sol.
What using it actually felt like
The recurring complaint is that K3 thinks too much and too slowly. Bindu Reddy cooled on the family because it spins and burns tokens; one tester watched a single SVG consume 33k tokens; another had a lone query time out after forty minutes at peak hours, against 28 tokens per second measured through one gateway. A separate strand of feedback calls the model overly proactive, pushing tasks forward before being asked. Against that, the capability demonstrations were striking: a coherent Minecraft scene built from spatial reasoning, a four-and-a-half-minute explainer animation of its own architecture, and a claim that K3 rewrote its own attention kernel over fifteen hours. Two behavioural observations are worth flagging as unresolved: K3's chain of thought resembles Claude's formatting closely enough that users suspect distilled training data, and the model depends on thinking history being preserved across turns.
The framing fight
The launch reopened the open-versus-closed argument. One widely shared reading is that the gap has closed and the "Chinese labs lag badly" framing needs retiring; others held that the promised weight release, not the benchmark table, is the part that would reset the field the way DeepSeek did. The sharpest counterpoint came on price: Moonshot has no Flash-tier product, which leaves it undercut by DeepSeek's cheaper line for the high-volume sub-agent work K3 is otherwise built for. One reposter raised the obvious downside of a capable open model — a lower barrier for custom builds aimed at fraud and intrusion.