AI News Daily · 2026-07-18
Today's summary
A model released the day before became the thing everyone else in the field was measured against, and the measuring got serious. Kimi K3 collected endorsements, benchmark places and viral demonstrations in the morning, then its first careful objections by the afternoon, most of them about what it costs to actually run rather than what it scores. Underneath that argument sat the question of who can afford to serve any frontier model at all: Meta was reported to be leasing capacity from a rival lab, Anthropic was reported to be arranging bank credit, and the price of server memory moved enough in a month to matter to everyone building. Away from the models, the day's most watched thing was not a model at all — it was a hall full of humanoid robots hitting each other.
- Kimi K3 spent its second day being audited rather than announced — Elon Musk called it remarkable at coding, a matched eight-task coding comparison against Opus 4.8, GPT-5.6 and Grok 4.5 placed it within reach of that group, and a "recreate macOS 27" instruction became the informal test everyone ran. One evaluation called it the strongest open-weight entry on LisanBench, and a forecast put it at 70.3 on ARC-AGI-2 if it holds up. The weights themselves are still a rumored date rather than a commitment, which did not stop a market opening on whether they are released at all.
- The objections were about total cost and reliability, not headline scores — Theo argued the interesting number is what a finished task costs end to end, because the model spends heavily on the way to its answers. Ethan Mollick's account of errors turning up inside a complex statistical audit travelled widely, one flat verdict placed the model below the Opus tier, and a separate observation noted the whole conversation is about capability with safety barely raised. Availability was patchy, which Miles Brundage attributed to a shortage of serving capacity rather than anything about the model itself. The distillation jibe, meanwhile, has decayed into a running joke rather than an accusation.
- The largest single report of the day put Meta on the buying side of Anthropic's compute — Meta is said to be negotiating to lease computing capacity from Anthropic in an arrangement potentially worth ten billion dollars, which is remarkable mostly for the direction of the transaction. In the same window Anthropic was reported to be arranging a multi-billion-dollar bank credit line, and its own availability notice prompted a question about whether the ceiling is demand or hardware. The imbalance that makes such a deal plausible was spelled out separately: one estimate has Anthropic holding fifty to a hundred and fifty times Moonshot's compute.
- The physical inputs tightened in the same week the demand grew — DDR5 server memory spot prices climbed 27.9% in thirty days, and one supply note described DRAM as the binding limit, with demand not the problem. Meta answered with land and people, committing to a roughly one-gigawatt Alberta campus and recruiting a senior AWS executive into its infrastructure organization, while Huawei showed a 950 SuperPoD rated at one exaflop. The dissent was loud and specific: Schmidhuber on what a trillion dollars of GPUs may fail to return, Gary Marcus on why commoditized models justify none of it.
- Open weights were argued as economics rather than engineering — The most widely shared framing of the release held that advanced capability is trending toward abundance instead of concentration, and Aravind Srinivas reached for the Linux-and-commodity-hardware precedent against Sun. Ethan Mollick noted how uncomfortably strong Chinese open-weight models have become, the White House AI adviser offered his own reading of the release, and the spread was quantified in a comparison putting a million frontier tokens anywhere from fifty cents to fifty-six dollars. Anthropic's premium drew a blunt list of pricing and moat questions.
- The practical rebuttal is that open weights are not the same as usable weights — A widely echoed complaint is that the recent wave of releases cannot be hosted on ordinary hardware, and the threads asking what it takes to run K3 at home read like a server build sheet. Compression work pushed the other way, with a 27B model quantized to 3.9GB and fitted onto an iPhone and a 120B mixture-of-experts streamed from storage on an Android handset. Thinking Machines' Inkling was named the highest-scoring open-weight result on ARC-AGI in the same window.
- The scaffolding around a model kept taking credit the model used to get — An Anthropic engineer's description of harness work as a rare and decisive craft circulated alongside a paper on optimizing that control layer directly, and the shift was named outright as loop engineering replacing prompt engineering. The plumbing arrived with it: Palantir launched a persistent execution layer that lets agents resume, Modal described running a million concurrent sandboxes, and durability was probed in discussions of surviving a crash mid-task and a benchmark for whether a fallback model inherits the whole conversation. Buyers were less enthusiastic, reporting that adoption stalls on evaluation, orchestration and hiring.
- OpenAI shipped steadily and gathered legal exposure at the same rate — The company reworked its new desktop application after user complaints, gave Codex a chat interface for pull requests, and had fresh specifics reported about its first hardware device. Against that it faces a new lawsuit tied to a suicide and, per the Financial Times, legal letters Apple has sent to individual employees — a dispute now prominent enough to be walked through in video roundups.
- Anthropic's day was mostly about how its own users feel — Complaints stacked up over a flagship model consuming extra credits, subscriptions cancelled over perceived degradation and refusals on ordinary questions. Product work continued underneath, with the Workbench rebuilt into a new Build area and Claude Code shipping a release carrying 48 command-line changes. An unconfirmed report has a new flagship arriving this week.
- Robots had the most watched moment of the day and, separately, the most useful one — China staged the first full-size humanoid fighting championship, whose clips escaped the field entirely and were called the best showcase embodied work has had lately. The substantive releases were quieter: Astribot's 4B robotics foundation model, Mimic's dexterous hand and wearable capture device, and an interchangeable gripper meant to let existing tools be driven by new software. The sober note came from a researcher on what has to line up before a failure shows itself in a real home, with a companion claim that the same simple method holds up across homes and offices.
- The revenue numbers arrived from the application layer, not the labs — Databricks disclosed an annualised run rate of $6.9 billion, up 80% year on year. Netflix told investors that roughly 300 film and television projects used generative tools this year. SAP completed its acquisition of Prior Labs, Intel is reported to be deploying Gemini Enterprise company-wide, and Zhipu is said to have passed a billion dollars in annualised revenue. The cautionary item is an account of a company paying $1.4 million a year for an assistant almost nobody opened.
- xAI kept shipping into the gap while generative video quietly improved — Musk teased Grok Imagine without detail, Grok added triggered, automatic task execution and a personalized timeline for model news, and Grok Build shipped terminal and session-jump improvements alongside users praising search over X from inside the tool. On the video side, Seedance 2.0 was used for finished cinematic sequences and testers said real-time character models are approaching livestreaming as whoever you like.
Since yesterday
- New: The day's most watched item has no counterpart yesterday — a full-size humanoid fighting championship that broke out of the field entirely. Also absent yesterday: the reported Meta and Anthropic compute arrangement, Anthropic's bank credit line talks, Astribot's Lumo-2 robotics model, Databricks disclosing $6.9 billion annualised, Netflix putting generative tools into about 300 projects, SAP closing the Prior Labs deal, and a first signal about Grok Imagine.
- Developing: Kimi K3 moved from launch to assessment inside a day. The endorsements landed — Musk on its coding and an open-weight record on LisanBench — at the same time as the pushback: the real cost of a finished task, mistakes inside a hard analysis, and the note that safety is barely part of the conversation, with the weights still only a rumored date. Grok Build, open-sourced yesterday, settled into ordinary iteration with terminal and session navigation and users demonstrating search over X inside it.
- Cooling: Yesterday's foundry results dropped out of view almost completely; that story survives today only as a note that equipment revenue is re-syncing with foundry capital spending. The dispute over a near-perfect harness score on the newest reasoning benchmark produced nothing further, and that family of evaluations appears now only as an open-weight leaderboard result and a predicted score for a model nobody has run on it yet. Talk of models improving their own training went quiet as well, reduced to a single system that automates frontier training runs.
coding & agent
Kimi K3's arrival reordered the day. Within hours of it becoming widely reachable, developers had wired it into OpenCode, Conductor and Grok Build, run it against the same eight-task suites they use for Opus and GPT-5.6, and turned a single prompt — recreate macOS 27 in a browser — into a shared informal benchmark. Underneath that noise, the more durable conversation was about everything around the model: the loop that drives it, the memory it keeps between runs, the sandbox it crashes inside, and the protocol layer that is starting to show its seams. The client tools moved too, in both directions — new capabilities in Claude Code and Codex, and a visible pile of desktop regressions filed against them.
Kimi K3 becomes the day's shared test bench
The fastest route to a verdict turned out to be a viral prompt. One developer ran a Kimi agent swarm at recreating macOS 27 with Liquid Glass effects, a run that took about three hours; another claimed Grok 4.5 did a comparable clone in under ten minutes, and a third noted that many observers rated the Kimi swarm above GPT-5.6 Sol Ultra on the same task. Structured comparisons were kinder to K3 than the theatre suggested: on one eight-task coding suite it matched or beat Opus 4.8, GPT-5.6 and Grok 4.5 on six of the seven tasks it finished. A Reddit thread asking whether K3 holds up in real codebases rather than benchmarks was the necessary counterweight.
What stood out more than the scores was stamina. One tester reported a run that hit the roughly one-million-token context limit twice and kept making progress through automatic compaction; another expected a quick prototype sketch and got a design exploration that ran past four hours. Distribution moved just as quickly — a documented path into OpenCode via a monthly subscription, three steps to reach it from Conductor, availability to OpenCode Go users before discount terms were settled, and a week of double credits. Output ranged from a Mac screen recorder submitted to Apple within a day to a Paper Mario-style game built with Grok Build.
The argument moves from prompts to loops
Several independent threads landed on the same framing: the leverage has shifted from wording to control flow. One widely shared piece argued that the past two years of agent improvement came from better prompts, and the next will come from better loops. Addy Osmani made the parallel case in a talk titled Own the outer loop, and a relayed account of an Anthropic engineer's view of the agent harness described that scaffolding as a rare engineering skill in its own right, decomposed into layers above the model. A paper pushed the idea further, breaking the harness into six editable control points that can be optimized directly.
The practical corollaries were unglamorous. LangChain argued that agents are not ordinary software and need their own development lifecycle, and scheduled a session on loop patterns for observable, testable systems. One engineer's advice on debugging was blunter: most failures live in the chain and the data interfaces, not the prompt, so add spans before touching the system message.
Memory and durability get treated as architecture
A Reddit account of a workspace that survived six model generations without agents losing memory or work state framed the question well: if the model is swappable, memory is the thing you actually own. Harrison Chase pushed wiki-style memory, so agents stop re-retrieving the same documents and rebuilding the same understanding, and separately relayed an open memory format proposed as a cross-agent standard. The mem0 founder made the strongest version of the claim, that continual learning is a memory problem rather than a training problem.
Durability is the other half. Palantir launched a persistent execution layer letting agents resume from their original state after crashes or waits, a Reddit thread canvassed approaches to surviving crashes in long-running agents, and Modal described scaling agent sandboxes to a million concurrent instances as the precondition for agents acting inside software at all.
Client tools gain features and shed stability
Claude Code shipped 2.1.212 with forty-eight CLI changes, including a /fork that copies a conversation into a background session; the release notes also cover session caps, background tool calls and a rename of the in-session subagent flow, while /code-review picked up selectable effort levels trading token cost against recall. Codex added PR Chat for reviewing and editing pull requests without leaving the current environment. Amp went further afield, letting an agent spawn another agent on a different machine and exchange messages and files with it.
The bug tracker tells the less flattering half. Codex Desktop on Windows drew reports of intermittent freezes after a recent build, hundreds of leftover taskkill.exe processes and CPU pinned near 100% at launch. Copilot CLI 1.0.71 was reported to leave unreaped zombie children on Linux and to hang on --resume at Windows cold start. OpenCode users said a forced new layout removed workspaces with no way back. One developer observed, separately, that Claude Code's issue tracker now carries more than eleven thousand open issues, much of it model-generated noise.
MCP's seams start to show
Adoption has outrun the safety story. One thread argued that connecting to an MCP server feels too much like piping curl to bash, because tool descriptions are themselves an injection surface, while another simply asked practitioners for their worst MCP security incidents. Builders of remote servers listed concrete traps, starting with OAuth 2.1 validation of externally issued tokens. Tool sprawl is the adjacent problem: Comet described cutting a server from thirty tools to four using self-correcting schemas and response compression, which sits against an open question about what happens when an agent is handed a thousand tools. Commercial questions are unsettled too, including whether per-call pricing survives agents that retry, explore and loop, and at least one shared piece argued MCP matters less than the discourse assumes.
Evidence from production, and the gap it exposes
The most concrete data point came from Anthropic's account of Bun's migration from Zig to Rust: roughly a million lines moved in under two weeks at around $165,000 in API cost. Coinbase described handing agents an entire internal pipeline in pursuit of self-improving loops. Against that, enterprise adoption keeps stalling on the same three things — evals, orchestration and talent — and procurement is shifting toward governable, measurable spending, while a critique of agentwashing noted that labelling ordinary workflows as agents has drained the word of comparative meaning. The skepticism has targets: one post contrasted Devin's claimed annual recurring revenue with an inability to find anyone actually using it.
Individual practice is converging on verification. One developer found that reading generated diffs aloud line by line dropped their acceptance rate to about 40%; another runs Claude and Codex as cross-reviewers of each other's work rather than as fallbacks. A database of agent failures now covering sixty traceable incidents found that roughly half involve security vulnerabilities. And there are new costs to carry: one developer reported a GitHub suspension after an agent submitted pull requests across their own repositories.
Apps
The product layer spent this window on plumbing rather than spectacle. OpenAI reshuffled the ChatGPT desktop client after user complaints, xAI turned Grok into something that runs on a schedule, and Google is reportedly preparing a place in Gemini for user-authored skills. Around those, the enterprise pitch hardened into named deployments, video tooling competed on asset libraries and project memory instead of raw generation quality, and a long tail of individuals shipped finished, installable software built in days. The counterweight, visible in the same window, is a run of reports about data leaking, quotas draining, and history disappearing.
Assistant clients get their housekeeping done
OpenAI rebuilt parts of the ChatGPT desktop app in response to feedback, pulling conversation history and Projects back into the sidebar and syncing chat and work history across web, mobile, and desktop. A smaller fix in the same spirit keeps the iOS dictation button visible when the input box already has text, with Android and web to follow. Both are the kind of change that only shows up when an assistant has become someone's daily tool rather than a demo.
xAI moved in a different direction, adding scheduled and triggered task execution to Grok so a described job runs on its own afterwards, and shipping a personalized timeline for tracking model releases. Its top tier now folds X Premium+ into SuperGrok Heavy for linked accounts, alongside the highest limits on Grok Imagine and Grok Build. Google, per a cited report, is building a native Skills menu into desktop Gemini where users could upload, create, and edit their own skills — the same portable-instruction pattern the coding tools normalized, arriving in a consumer chat surface.
The enterprise pitch stops being abstract
OpenAI put out customer material rather than claims: Shopify's Chris Jones described restructuring internal processes around ChatGPT Work by assigning discrete tasks to agents and leaning less on engineering, and a Virgin Atlantic team walked through summarizing strategy documents and competitor analysis. Google's CEO said Intel is adopting Gemini Enterprise company-wide, including in semiconductor R&D. Elicit described a year of continuous use at Orion Pharma, starting with reports and systematic literature reviews before widening.
The vertical entries kept coming: AWS introduced Amazon Quick, an agentic assistant for sales teams that reads CRM, email, Slack, and tickets, and Microsoft moved Copilot Health into preview. Notably, the same work-branded product is being pulled sideways into private use — one user for personal email, calendar, and household admin, another for tracking workouts and maintaining a fitness plan.
Creative tools compete on context, not generation
HeyGen spent the window on supply rather than models, attaching a large media library to HyperFrames with music, images, sound effects, and logos free to logged-in users, then adding word-timed animated caption styles. TapNow pitched itself as a creative operating system that keeps references, methods, and project context resident in the workspace long-term — the explicit argument being that persistence, not output quality, is the bottleneck. Invideo's Agent One was reported to build a music video through sustained conversation, and Uisato showed a Music Video Pro mode accepting multiple image references.
Around that, practitioners traded economics and craft: a hands-on test of CapCut's Director Mode for character consistency across unrelated scenes, a trick for cutting the credit cost of video extension in Dreamina, and Seedream 5.0 arriving on OpenArt with usage uncapped for now. Higgsfield opened a free filmmaking course. One limit stayed stubborn: a user found that recoloring part of an image without disturbing its text defeated several mainstream tools.
Individuals shipping finished software
The build reports this window were about installable products, not prototypes. A developer used Claude Code to make a text-to-speech app handling links, PDFs, and photos; another replicated a paid screen recorder as a native Mac app in under a day with Kimi K3 and submitted it for App Store review. Open-source entries included a canvas that puts Claude's answers beside handwritten notes, a local dual-brain desktop agent bundling Ollama, a screen-watching monitor that alerts through messaging apps, a keyless AI browser, and a Windows PDF suite aimed at Acrobat's workflow.
Games were the other pocket of energy — a jam entry built with the generation tool Spawn with complete loops, and a board game made with Vibe, now opened to the public after starting as an internal tool. Roblox, meanwhile, said users will soon create games from prompts on-platform.
Connectors spread, and so do the failures
The integration surface widened. OpusClip exposed its video editing as an MCP interface callable by agents; a developer released a webhook tunnel built for agent debugging to replace ngrok-and-paste; Grok Build gained direct X search inside the build flow. A travel operator tested whether an unfamiliar agent could actually find their MCP endpoint, which is the discovery problem the registries have not solved. Unabyss went further upstream, consolidating scattered personal context into one vault for tools to draw on.
The failure reports landed in the same window and are worth equal weight. A Reddit poster said Prism returned other users' papers in compiled results, a claimed isolation breach. Another reported Gemini Pro deducting quota for prompts that never returned. ComfyUI's desktop 28.0 update brought memory errors to workflows that previously ran, and a Windows reinstall of Claude Desktop reportedly wiped local Code session history while account-synced chats survived.
Research
Robotics and world models took up most of the oxygen, but the sharper movement was in what researchers will accept as evidence. Several of the strongest arguments were not about new capability at all — they concerned the conditions under which a capability claim should be believed: an unfamiliar home rather than a demo table, a neutral evaluation rather than a launch-day table, a replication rather than one lab's plot. Underneath, architecture and efficiency work kept migrating out of preprints and into the stacks frontier labs ship, while AI-for-science produced both a striking claim and pointed skepticism about itself.
The unseen-home standard for robots
Tony Zhao laid out the tightest statement of the day on robot evaluation: a problem only surfaces when three conditions hold at once — the system runs fully autonomously, it is tested in a real home it has never seen, and the policy is not quietly assisted. He argued that dropping any one of the three hides exactly the failures worth finding. It sits awkwardly beside the same author's relay of a more excitable claim that general robot intelligence is close, and beside a longer-running argument that laundry folding will be the first home task solved because current methods happen to suit it.
The supply side moved in parallel. Astribot released Lumo-2, a four-billion-parameter robotics foundation model that predicts task-relevant future physical states before choosing actions, with the author noting it does not generate full future video. Mimic Robotics announced the M1 dexterous hand and the U1 wearable as parts of a full physical-AI stack. On measurement, RoboDojo compared thirty policies across forty-two simulated and eighteen real manipulation tasks, and separate work on long-horizon control addressed why naively extending visuo-motor context blows up inference time.
World models, still contested at the base
The gap between "world model" as a research program and as a marketing word stayed wide. A YC discussion framed the core problem plainly — why frontier systems still need thousands of samples for skills humans pick up in a handful — while a hiring note signalled that the VL-JEPA lineage is being staffed up. Against the simulator framing, one widely shared test proposed asking a video model to predict balls colliding one after another, on the argument that sequential dependency is where standard video generation breaks.
Concrete systems came from several directions: a latent world model built on SIGReg that its promoters describe as cheap to train and competitive with DINO-WM; DeepMind's demonstration of turning raw walkthrough video into a scene queryable in natural language with no manual annotation; and Alibaba DAMO's teleoperation work, which replaces the remote-controlled robot with human gestures driving a generative world model. A counterweight arrived from adversarial research: BadWAM argues that coupling action generation to future prediction does not confer the robustness people assume, since a model can predict correctly and still act wrongly.
Science moves from promise to specific claims
The window's loudest science claim was that a system reached independently, in about two days, a conclusion a lab needed roughly a decade to confirm — reported without the underlying detail, and worth treating as an assertion until it is. Around it sat the machinery that would make such things routine: AutoScientist for automating frontier training aimed at discovery, an agentic system that runs literature search through random-effects meta-analysis end to end, and a defended thesis asking directly whether language models can participate in open-ended reasoning and discovery. A podcast thesis put the strategic case bluntly — with the internet largely consumed, scientific data is the next corpus at internet scale — and one researcher listed four priorities for funders and policymakers, starting with making agents widely available.
Biology supplied both the applications and the doubts. An all-atom diffusion model targets macrocycle conformation prediction and design, and an autonomous lab published results from over a thousand solid-state synthesis experiments. But a long piece on AI-designed antibodies concluded the science is solid while the business case remains genuinely uncertain, and another summarised the field's position as "demo or die" — atlases everywhere and no destination.
Efficiency work lands inside the frontier stack
The Kimi line drew unusual attention. Its delta attention design was described as an eighteen-month iteration begun in January 2025, with an independent blog breaking down the trade-offs against recent alternatives, and an Nvidia paper circulating for proposing a modified latent-MoE variant of the 1T model. Tying the thread together, hybrid linear attention, kernel benchmarks and full-model megakernels are now applied directly in frontier model construction rather than left academic, and Moonshot's own GTC talk covered how the team keeps approaching closed-source frontiers with open models.
Method-level work followed the same compression instinct. DeepLoop proposes looped transformers as a depth-scaling research vehicle; Latent Thought Flows squeezes 256 text tokens into eight continuous latents before decoding back; a reinforcement learning recipe simply penalises reasoning tokens inside the reward; and LongStraw targets million-token post-training within a fixed GPU budget. On distillation, one paper catalogued the roles and failure modes of on-policy distillation, a useful corrective to its current popularity.
Trust, but verify the verifier
Evaluation scepticism was the connective tissue. A new translation benchmark was built specifically because existing ones are saturated and automated metrics are fragile. One widely echoed warning was to treat vendor-published launch scores as marketing and prefer neutral or self-run evaluations. Another argued that localised evals over-optimise the wrong things without an end-to-end view, and a compact formulation held that metrics can be aligned, unbiased and simple — pick two. Contamination got its own reminder: text detectors cannot be validated on pre-ChatGPT public texts the models have already read.
Interpretability moved the same way, toward replication. One researcher reproduced Anthropic's J-space analysis and extended it with new experiments, while another noted a methodological caveat that using 250 prompts instead of 1000 changes the estimated shape. A separate thread analysed evidence that hidden states linearly encode whether a model will comply before it emits a token.
Security research reached a structural verdict: an arXiv position paper argues agent security is systemic and unfixable by point defenses, while a hands-on experiment put five agents — triage, development, scanning, review, deployment — in a CI/CD pipeline and fed them a single malicious ticket as untrusted input.
Models
Moonshot AI's Kimi K3 landed just before this window opened, and for the next twenty-four hours almost every other model story was told in relation to it. The release is described as a frontier-class open-weights model with a very large parameter count, a million-token context and a native multimodal architecture, and the response split quickly into three arguments that ran all day: whether the benchmark placement survives contact with real work, whether the price advantage holds once you count tokens rather than rates, and whether "open weights" means much when almost nobody can host the model. Around that, OpenAI's GPT-5.6 line kept a reputation for steadiness, Anthropic spent the window handling a billing fault and a wave of refusal complaints, and several smaller open releases went out with little notice.
A release that reset the frame
The framing that spread furthest was not a score but a claim about supply: that advanced AI now looks more likely to be abundant than scarce, and that export restrictions on frontier compute do not obviously change that trajectory, argued by Jeremy Howard. Coverage picked up the same theme from the other end, noting that a team of roughly 300 people produced something early reviewers put near Opus 4.8, with a fuller specification writeup in the AI News issue. Endorsements arrived from unusual places: Elon Musk called the model's coding genuinely strong and its design sense good while noting spelling errors, and US AI czar David Sacks weighed in publicly enough to become its own story. The pushback was equally prompt: Seán Ó hÉigeartaigh rejected the "China has closed the lead" reading while granting that K3 is impressive, and Ethan Mollick noted that the ranking among Chinese open-weight models is shifting fast. Weights themselves are not out; a speculative post put the date at July 27 at even odds. Moonshot's own account of how it got here came via Yang Zhilin's GTC talk on token efficiency and open-source scaling.
What the tests actually showed
On evaluations the picture is consistent and bounded: strongest open-weight entry, not the strongest entry. K3 leads open weights on LisanBench, ahead of Gemini 3.1 Pro but behind Opus at high effort. An eight-task side-by-side against Opus 4.8, GPT-5.6 and Grok 4.5 had it matching or beating all three on most tasks it completed, and a screenshot put it first on a Next.js evaluation without methodology. Demonstrations were the loudest evidence: a 48-hour autonomous chip design run using open EDA tools, and a prototyping session that ran continuously for over four hours.
The counter-evidence is narrower but pointed. Mollick reported multiple statistical errors during a complex audit, including misapplied methods. A blunter assessment placed it short of Opus tier on long context and multi-turn agentic loops, and a developer warned that under its maximum reasoning setting the model is poorly matched to basic tasks — a caveat that should qualify any evaluation. Practitioners on Reddit spent the window asking for first-hand results in real codebases rather than scores.
The cheapness claim, examined
The most useful correction of the day came from Theo, who pointed out that K3's per-token price is half GPT-5.6's but its total real-world cost lands in roughly the same place because it emits far more tokens. A related post credited OpenAI with an inference-efficiency advantage measured in output density rather than headline rates. Subscriptions invert the comparison again: cheaper API pricing can work out more expensive inside a plan. Supply is the other constraint — the service was visibly straining under load, one user reported it degraded within a day of launch, and the natural reply was release the weights if you cannot serve it. Tooling moved anyway, with opencode offering double usage credits for a week.
Open weights meet the hardware floor
A recurring complaint is that recent open releases are not locally runnable in any ordinary sense, licences and context lengths notwithstanding. Estimates for K3 circulated at four 512GB Mac Studios, and a separate post worked through the memory arithmetic for a model at this scale. Hence the obvious ask: distil it into a flash-tier model. Meanwhile the small end kept delivering — an aggressively quantized 27B fitted onto a single 3090, DeepSeek V4 Flash reached a million tokens of context on a 5090, and a heavily quantized MacBook came close to a pair of DGX Sparks. Other open releases passed with little attention: Inkling took the top open-weight ARC-AGI scores, Tencent Hunyuan put out a 295B mixture-of-experts model, and NVIDIA shipped a new retrieval embedding model.
Closed labs under pressure
GPT-5.6 held its ground on quality — Sol Pro nearly maxed one private benchmark and the line was called more reliable than its Anthropic counterpart — but at a cost, with users hitting capacity limits on consecutive days and reports that full-access mode deleted an entire home directory. Anthropic had the harder day: it acknowledged overcharging on extra usage and promised refunds plus compensation after users noticed credits drawn down early, alongside cancellations over refusals and error rates and a Hacker News thread about a Kindle project blocked as suspicious. The competitive read is that release timing is being pulled forward, with Opus 5 rumoured for this week and a claim that K3 is accelerating frontier cadence generally.
Multimodal
Two shifts ran through the day's multimodal material, and they point the same direction. Latency kept falling until generated video started behaving like a live medium rather than a render job, and the editing decisions — which take, which cut, which reframe — kept moving from the person to the model. Around both, the open-weight side filled in the unglamorous parts: video understanding, audio restoration, 3D completion, and a batch of benchmarks built to measure whether any of this actually holds together.
Generation that answers in real time
The clearest single data point came from hands-on testing of Lucy 2.5, where the reported latency was low enough that a tester framed the capability as livestreaming as any character. A separate prototype built on Omni went at the same idea from the interaction side, demonstrating directly manipulable real-time video as a proof of concept rather than a shipping product. Both are early, and both were presented as such.
The infrastructure underneath moved in the same window. Wan-AI restructured its streaming interaction model around a framing of video as a stable world plus a time-varying event stream, while Tongyi Lab's release notes for the preceding version described a single end-to-end transformer meant to listen, see, speak and act at once. One developer added speech to a live image generation system so it listens and answers while drawing. On the closed side there was mostly signalling: Elon Musk posted about Grok Imagine without detail, and a circulating roadmap said to be Grok 4.1's lists near real-time generation among its features — an unverified leak, and worth treating as one.
The model takes the editing chair
Polymarket relayed a claim that Kimi K3 assembled a trailer from 56 clips unaided, work described as one to two days for a senior editor. Runway published an independent evaluation of video agents scored across thirty cinematic prompts and sixteen metrics, in which its own agent ranked first on narrative coherence, cinematic language and production quality — a vendor-shared result about a vendor's product, though the methodology is at least stated. Elsewhere the same pattern showed up in smaller form: a user turned a personal song into a music video by talking an agent through revisions, and another handed nine product photos and a rough brief to Claude Code and came back to a finished cut.
Practitioners are meanwhile building the connective tissue by hand. fal shared a restyling chain that uses depth to hold motion and composition while swapping the look, and another workflow paired generation with automatic reframing instead of treating a single shot as the deliverable. HeyGen fed the same appetite by opening a large stock media library to logged-in users.
The counterweight was loud. One widely shared argument held that generating an explainer video is easy and knowing what makes one good is the hard part. Cost surfaced repeatedly, with one creator publicly hunting for a cheaper way to produce short close-up clips. And quality is uneven: a tester called LTX 2.3 close to unusable for image-to-video, citing faces distorting within seconds.
Spatial output as a first-class result
Several threads converged on models emitting geometry rather than pixels. Moonshot positioned Kimi K3 around 3D reasoning fed by vision, turning images and video into interactive scenes; a builder used Fable 5 with three.js to make stadium seat-view previews in five prompts, and another reported building three navigable 3D worlds with GPT-5.6 Sol.
Reconstruction advanced alongside generation. SpAItial's Echo infers and fills the holes that real-world scans leave behind, an open pipeline based on Trellis.cpp reached reference-implementation asset quality, and BRDFusion decomposes street video into geometry, materials and lighting. Google's open-sourced GNM offers a structured way to describe a character's body and face, already demonstrated with Blender vertex-group masking. Gaussian splatting supplied the day's rendering tricks, including convincing fake reflections in WebGPU.
Open weights, and something to measure them with
The open releases spread across every modality. VideoChat3 arrived as a fully open video understanding model with a paper behind it, SenseNova U1 was described as an underrated open image model and its infographic variant added local text correction on already-generated output. On audio, Diamond 1.0 restores degraded speech to studio quality and the Wan team's report describes five-minute songs with separated vocal and instrumental stems. NVIDIA extended its fine-tuning library to Hugging Face Diffusers, which pulls image and video models into the same tuning path as text.
Evaluation showed up in unusual volume. The Kling team put out benchmarks for both keyframe-conditioned video and multi-reference audio-visual generation, a university group released a video benchmark shot from the first-person view of visually impaired users, and ByteDance's UniVR probed what a model can learn from purely visual demonstration. A sharper test came from a skeptic, who argued current video models fail at predicting a chain of sequential collisions — a small ask that separates a convincing renderer from a world model.
Infra
Compute started behaving less like a capital asset and more like a traded commodity. The window's largest story was a reported ten-billion-dollar lease negotiation between two rivals; its most consequential may be memory, repricing fast enough to constrain everything built on top of it. Around those, China's domestic stack consolidated around Huawei, the serving stack shipped a run of unglamorous throughput work, and tinkerers kept pushing frontier-scale weights onto hardware that has no business running them.
Compute becomes something you rent from a competitor
The window's dominant item was a report that Meta and Anthropic are discussing a compute lease worth as much as ten billion dollars, attributed to The New York Times in a second account. Accounts differ on which way the capacity flows: one framing has Meta renting out idle capacity with Anthropic as an early customer. Either direction points the same way — the firms building the largest fleets will now sell access to them, and buy it from people they compete with. Meta's other moves fit that reading: a first Canadian data center in Alberta at roughly 1GW and about $9.2 billion, and the hiring of a senior AWS executive onto its compute team, per WSJ.
Financing is following the same logic. General Compute took up to $400 million in debt from Upper90 to buy ASICs, a deal read elsewhere as inference chips serving as collateral and as a shift in how AI hardware gets financed — away from betting that training produces a model, toward betting that inference produces revenue. Skeptics were audible: Schmidhuber warned of an infrastructure trap in which a trillion in GPUs becomes a $900 billion write-down, Gary Marcus asked why trillions are still going into data centers as models commoditize, and one widely-shared thread predicted big labs will resell what they cannot consume to long-tail enterprises.
Memory is the binding constraint
DDR5 server RDIMM spot prices rose 27.9% in thirty days to $1,375, and SK Hynix's $26.5 billion ADR offering was reportedly oversubscribed roughly sevenfold — read as a market bet on a durable HBM shortage. One infrastructure vendor described DRAM, not demand, as the ceiling on its own upside. The squeeze is already reaching consumers, with the AI build-out named as a factor in India's smartphone memory crunch.
Elsewhere, US chip imports from Taiwan hit $1.49 billion a month, up 64.4%; equipment vendor revenue is re-syncing with TSMC's capex after decoupling in 2023-24; and voltage regulators were flagged as an overlooked choke point, with Monolithic Power Systems said to hold about 70% of the slots on NVIDIA's newest GPUs. Beth Kindig's summary of where the money went is blunt: Micron has added more market value this year than big tech combined.
China's stack consolidates around Huawei
Huawei introduced the 950 SuperPoD, claiming 1 EFLOPS fp8, 2 EFLOPS fp4 and 256TB of unified memory, alongside circulating technical analysis of its Ascend system design. One reading of the K3 release — which uses MXFP4 rather than NVFP4 — is that Chinese models may soon tune specifically for Huawei silicon. The software layer moved in parallel: Metax pitched its open-sourced GPU stack as an Android for domestic accelerators, Biren said it had adapted its stack to GLM-5.2, and Zhipu acquired compiler shop Zhongke Jiahe outright. The framing through the Chinese coverage is that the contest is no longer peak FLOPS but deployable system capability and end-to-end token production. Constraints remain real — Miles Brundage restated that China has far less total compute — and gray-market workarounds persist, including reports of RTX 5090s stripped and rebuilt into 128GB server cards.
Throughput work in the serving stack
The efficiency news was cumulative rather than dramatic. SGLang reported 6.4x higher throughput by reusing KV caches for shared prefixes via a radix tree, while vLLM with LMCache showed 2.8x on repeated prompts with no extra GPUs — two takes on the same insight, and one practitioner called KV cache management among the hardest problems in inference systems. Releases landed across the stack: PyTorch 2.13 with 3,328 commits, ROCm 7.14 with a new open build system, MXFP8 support in prime-rl, RL post-training framework Vime running natively on AMD Instinct, and NVIDIA's Nemotron-3-Embed-8B retrieval model.
Operational realism showed up too: arguments that production pain lives in tail latency, not medians, that structured generation must satisfy performance, correctness and schema adherence at once, and that semantic caching's apparent hit rate hides false reuse of merely similar queries. Efficiency is also becoming the headline metric — NVIDIA published findings on power-constrained system design and pitched Vera Rubin on intelligence per dollar, while one tinkerer measured tokens per joule on a power-limited P100.
Frontier weights on hardware that should not hold them
The local scene spent the window testing how far compression goes. PrismML's Bonsai, a 1-bit version of a 27B Qwen model, was reported shrunk from 54GB to 3.9GB and loadable on a single 3090 without retraining, with tests suggesting 1-bit weights are faster as well as smaller. Applying the same treatment to Kimi K3 is already being floated as a path to single-machine deployment, against unquantized estimates of four 512GB Mac Studios.
Hands-on results were mixed but striking: a heavily quantized MacBook running DeepSeek-V4-Flash came close to two DGX Sparks on a terminal benchmark, another user pushed the same model to a million tokens of context on a 5090, and a 120B mixture-of-experts model was streamed off flash storage on an Android phone at 1-5 tok/s. Sparsity is why any of this works under bandwidth limits, as side-by-side local measurements showed. At the small end, a Google engineer reported fine-tuning Gemma 270M on a phone in 21 minutes, lifting task accuracy from 46% to 90%.
Embodied
Two very different pictures of embodied AI shared the window. One was a spectacle: a full-size humanoid fighting championship in China, where a robot kicked its opponent's head off and then danced. The other was quieter and more consequential — a wave of WAIC demonstrations built around long, unscripted tasks, a run of new robot foundation models and benchmarks, and a hardware layer arguing about hands, motors and where the compute sits. Underneath both, practitioners spent the day arguing about what "solved" should even mean.
The fighting league and the WAIC floor
China hosted what was billed as the world's first full-size humanoid fighting championship, and one of the accounts carrying it was explicit that the footage spreads because it looks like a stunt rather than because it demonstrates capability, noting the head-kick and the dance routine in the same breath. Others treated it more warmly, calling it one of the more entertaining showcases embodied intelligence has produced lately, and noting that the leagues themselves are scaling up.
The WAIC material pointed the other way. Yuanli Lingji ran a public challenge in which six robots worked continuously for fifteen hours to build a wall out of 80,000 blocks with no teleoperation and no pre-written script. Qianxun Intelligence demonstrated continuous long-horizon work such as tidying a living room, rather than isolated grasps. Pudu Robotics used the show to argue a "one brain, many forms" data strategy, and Alibaba's DAMO Academy proposed replacing physical teleoperation with human gestures driving a generative world model that synthesizes first-person robot video. A separate account of the first PhysicalAI hackathon in Haidian read the supply chain and deployment capability behind that ecosystem rather than the event itself.
Models, benchmarks and their failure modes
Astribot released Lumo-2, a 4B robotics foundation model that predicts task-relevant future physical states before choosing actions — with the caveat, from the person describing it, that it is not generating full future video. RxBrain proposed an embodied cognition model that runs language reasoning and visual imagination inside a single planning sequence, and a paper circulated on feeding robot-centric pointmaps into vision-language-action systems so that varied camera angles stop making learning harder. Another thread covered work on scaling visuo-motor context to long horizons without inference time blowing up.
The counterweight arrived the same day. RoboDojo compared thirty representative policies across 42 simulated and 18 real robot tasks under one standard, and BadWAM argued that world-action models, often assumed safer because they predict the future, can be pushed into thinking right and acting wrong. A penetration-free cloth simulation dataset of roughly 493K frames was released, and a YC episode framed the open question plainly: why frontier systems still need thousands of samples for skills people pick up in a few tries.
Hands, motors and onboard brains
Mimic Robotics announced the M1 dexterous hand and U1 wearable, pitching a full physical-AI stack. A competing view shipped alongside it: an interchangeable gripper whose makers argue humanoid hands fall short where stability matters most. One long piece made the case that humanoid robots must carry their brains onboard, since interaction happens in milliseconds and cloud round-trips do not; another treated axial flux motors as the foundational capability for robots and vehicles alike, and a third laid out a Sys0-Sys1-Sys2 hierarchy with kilohertz low-level control at the base.
On the business side, Agility Robotics is opening a Digit training center in Fremont, NVIDIA is extending its physical AI ecosystem in Japan, Ontic Labs launched and started hiring after selection into a European program, and unconfirmed market chatter had Figure raising again at a reported valuation the original poster did not stand behind.
What counts as solved, and who pays
The most useful argument of the day was definitional. One developer proposed that a task is solved only when users will pay for it, not when a demo succeeds; another set three conditions for a test to surface real problems — full autonomy, an unseen real home, and a fixed policy — and a third revisited the case that laundry folding suits current methods well enough to fall first.
The consumer edge moved too. A report described OpenAI's first device as a screenless, battery-powered speaker with moving parts meant to look alive; Qwen upgraded its glasses to Agent Glasses and added earbuds; Moonix shipped a lightweight standard edition; and one essay argued glasses should be built as AI's perceptual memory organ rather than a phone replacement. The friction is real as well: workers at a Hyundai plant went on strike over humanoid robots arriving on the factory floor.
Venture
Money moved in two directions over the window. Operating companies published revenue figures large enough to make the sector look self-funding, while the people writing the checks spent the day arguing about valuation discipline and about where the cash for compute is actually coming from. Primary deal flow was modest and scattered: SAP completed its acquisition of Prior Labs, which says it will keep operating as an independent lab; agent company Lyzr reported a $100 million Series B at a valuation near $500 million; and audit-automation startup Auxilius came out with €1.3 million.
Revenue milestones and the gap between claimed and observed
Databricks said it is running at $6.9 billion in annual recurring revenue, up 80% year over year, with $1.7 billion of that attributed to AI products. On the Chinese side, Zhipu is reported at $1 billion ARR — sourced to local media rather than the company, so it belongs in the claim column for now.
The counterweight landed the same day. Cognition's Devin was said to have passed $492 million in ARR, a number one poster set against being unable to find a single person actually using the product. Whether or not that is fair, it captures the season's recurring problem: a revenue line no outsider can reconcile with what they observe. The base rate underneath all of this is harsher still — a tally of 8,281 startups found just over half at zero revenue and another third under $1,000 a month.
The bill is increasingly financed rather than paid
Anthropic is negotiating a multi-billion-dollar credit line with banks ahead of a planned listing, according to The Information — liquidity assembled before an IPO rather than raised by one. The New York Times made the general version of the argument, reporting that AI spending leans more and more on borrowed money. Hardware is following: a $400 million loan collateralized by chips shows lenders rotating from training hardware toward inference silicon as the asset they will lend against.
Compute is also drifting toward being traded rather than owned. One investor sized a potential US compute futures market by comparing derivatives volume to annual production across other commodities, and Meta is reported to be in talks with Anthropic on an arrangement worth up to $10 billion, described elsewhere as Meta renting out idle data-center capacity with Anthropic as an early customer. Public money took a different route: SPRIND named ten teams from five countries eligible for up to €26.5 million each, non-dilutive, and one participant argued the real value of equity-free money is the private rounds it makes easier rather than the ownership it preserves.
Margin pressure meets valuation nerves
Benchmark partner Everett Randle warned that investors have settled into a blind conviction that every good AI company will be worth more in six months, which he reads as a 2021 echo. Gary Marcus went at a specific target, questioning xAI's valuation on the grounds that Musk's advantage in AI is thinner than the price implies.
The structural worry is margins. One widely shared note argued that inference costs are compressing gross margins at AI application companies, setting up a re-rating in growth-stage private markets; another held that the harder frontier labs push for high margins now, the worse the eventual reckoning once leaner competitors arrive. The price gap between US and Chinese inference is the mechanism. Meanwhile Apple, written off as an AI laggard, passed Nvidia in market value intraday for the first time since last spring, and unverified chatter had Figure raising again at $50 billion to $80 billion.
Safety
The day's safety conversation split cleanly into two halves that rarely talk to each other. One half is about harm to people: a fresh wrongful-death suit against OpenAI, Meta wiring parental alerts into teen conversations, and San Francisco prosecutors going after apps that undress real photographs. The other half is about harm through machines: agents that approve malicious code, tool descriptions that double as injection vectors, and a model that wiped a user's home directory. Regulators moved on both fronts, but unevenly — Europe and a handful of smaller states shipped concrete instruments while the main US technical authority was described as boxed in.
Liability lands on chatbots and app stores
OpenAI is facing another suicide-related lawsuit, this one brought over Christian Faith Madison, who is said to have died after months of conversation with ChatGPT that allegedly reinforced her framing of what she was doing. The claims are untested, but they arrive alongside a psychiatrist's argument, circulated by Zak Kohane, that long chatbot sessions themselves carry measurable psychological risk. Meta's response to the same pressure was product-side: teen conversations with Meta AI that touch self-harm can now trigger a notification to parents.
Enforcement showed up in the app stores. The San Francisco City Attorney sent cease-and-desist letters to Apple and Google over thirteen face-swap and "nudify" apps, demanding they stop profiting from them and asking for removal outright. TikTok, meanwhile, began testing an opt-in tool that lets verified creators scan for synthetic fakes of their own likeness and report them — consent infrastructure built voluntarily, ahead of any rule requiring it.
Rules arrive, but the referee is weak
The most concrete regulatory action came from Brussels: the EU introduced requirements that Google share search data and open its AI ecosystem on Android, a demand echoed in US reporting that Google is being pushed to give AI rivals room on the phone. Estonia went further into new territory, planning to be the first country to issue identity codes to AI agents so the state can verify on whose behalf an agent is acting. Australia attached an energy condition to data centers, requiring them to generate more power than they consume, and twenty-nine countries signed on to a global AI cooperation body.
US action stayed at the state level and stayed contested. Miles Brundage argued that xAI's Grok 4.5 may already breach California's SB 53 given how quickly it was jailbroken, and pointed to the asymmetry with rivals that published lengthy safety documentation. Others defended Massachusetts S.2630's catastrophic-risk assessment and independent review clauses and pushed back on the framing that such rules hurt startups, noting the thresholds target the largest and highest-revenue firms. Against all this, a report on the Center for AI Standards and Innovation described an office that could oversee frontier development but is constrained in practice.
Agents are the new attack surface
The security material converged on one claim: point defenses will not hold. An arXiv paper argued agent security is systemic, failing at tool calls, permission boundaries and context passing rather than at any single gate. A five-agent CI/CD experiment gave the pipeline one malicious ticket and watched the reviewers wave the resulting code through. Practitioners compared connecting to an MCP server to piping curl into a shell and asked what an allow-list audit should even cover, while a separate thread collected real MCP security incidents.
Concrete exploits kept pace. A macOS Terminal chain paired indirect prompt injection with DNS-based exfiltration, Iranian operators ran phishing through AI-generated LinkedIn recruiter profiles, and ransomware is reportedly turning agentic. GPT-5.6 in full-access mode was found to delete an entire home directory in multiple reported cases. The containment answers were sandboxes and interceptors: Perplexity's SPACE isolation platform and a guard layer for OpenClaw agents that gates high-risk actions.
What we still cannot measure
Underneath both halves sits a measurement gap. Brundage pressed US labs to disclose more about safety post-training, arguing the case has grown stronger with recent releases — a point sharpened by complaints that debate around Kimi is all capability and no guardrails. SecureBio filled one slice of that gap with BioTIER, presented as the first benchmark aimed squarely at biosafety refusal behavior.
Calibration cuts the other way too. A developer trying to turn an old Kindle into an e-ink monitor had the request refused as jailbreaking by two frontier models, the kind of false positive that erodes trust in guardrails generally. And one researcher noted that even a published model constitution can encode the vendor's own interests in ways users cannot audit.
AGI Musings
The commentary in this window kept returning to one premise: that frontier capability is becoming cheap, and that most of the industry's assumptions were built for a world where it was not. A fresh open-weight release from Kimi supplied the occasion, and the arguments fanned out from there — whether scale is a durable advantage, whether compute rather than model quality now decides who wins, what happens to software once agents sit between people and machines, and what is left that is distinctly human once answers cost nothing. Little of it was new argument. What was new was how confidently people held their positions.
Abundance versus scarcity
Jeremy Howard put the case plainly: advanced AI is likelier to end up abundant rather than scarce, and restrictions on US frontier models and compute do not obviously change that. Aravind Srinivas reached for a historical parallel, comparing the moment to Sun Microsystems losing to Linux, x86 and commodity hardware. A blunter version circulated as the "frontier curse": the better a closed model performs in your industry, the more urgent your need to control your own weights.
If the moat is not mysterious technology, what is it? One argument on Reddit held that Anthropic's and OpenAI's advantages reduce mostly to scale, with a companion post asking whether a 27B open model reaches today's frontier within five months. Gary Marcus pressed the same point from the money side, questioning trillion-dollar data center commitments if models keep getting smaller, more open and more efficient. The margin version of the worry: labs pushing hardest for high margins now may look worst once leaner competitors land. Dean Ball offered the ecumenical read — the best outcome has room for both open and closed — while others worried about a subtler cost of a distillation-heavy world, in which models converge on one personality and lose their distinct voices.
Compute as the deciding variable
Running against the abundance thesis was a compute-first argument. Peter Wildeford's version: even if China and the US hold equally capable models, whoever mobilizes more compute deploys more AI and takes the advantage. Jensen Huang's framing was cited to the same end — agentic systems plan, call tools and iterate, so they push demand to a different order than single-pass answering. A more concrete estimate asked what happens when hundreds of millions of robots each run a trillion-parameter model in real time.
The geography of that spend drew its own commentary: China's capital expenditure remains below Western frontier labs but is closing, Europe is structurally disadvantaged on funding, compute and salaries, and one argument held that compute limits could cap Chinese labs at the domestic market. Domestically, the same buildout was read as a concentration of wealth around whoever owns the power, land and silicon.
Agents, interfaces and work
Elon Musk's claim that browsers, websites and traditional software all disappear in favour of direct agent use is the maximal form of a shift people report anecdotally: one Reddit user described using ChatGPT as an ad-free replacement for browsing. At the research end, a16z's David George described teams moving off the keyboard entirely, directing swarms of agents by voice.
The counterweight was empirical. Anthropic's own numbers were read as the gap that matters: AI touches roughly 60% of engineering work, but under 20% ships unreviewed. One critique argued that what is being sold as continuously learning agents is closer to a frozen state machine, and another that narrow evaluations over-optimize the wrong things without an end-to-end view. The labour consequences surfaced concretely — Hyundai workers struck over humanoid robots on the factory floor — and speculatively, in the suggestion that white-collar roles get platformized and fragmented the way delivery work was. Tyler Cowen took the other side, arguing the era creates meaningful new work rather than only removing it.
What stays human
The consciousness argument reignited, with functionalism set against quantum accounts of awareness, and a cleaner reply offered by analogy: flight arrives by flapping or by jet, so coherent language need not imply consciousness. Whatever the answer, one view held that the argument itself drags moral consideration outward.
Nearer to daily use, two separate threads made the same observation — that the burden of being right has moved to the reader, because a model produces plausible hypotheses rather than sourced answers, and that accuracy is now the user's job rather than the author's. A practitioner's complaint sharpened it: high error rates in specialist domains are masked by fluent, confident phrasing. Alongside came the losses that are harder to measure — that tools cannot teach taste, and that shipping more while doing less demanding work costs a sense of mastery. Hence the recurring case that this is exactly the wrong moment to cut philosophy teaching.
Companies & People
Corporate news in this window sorted into a few hard lines. Apple's dispute with OpenAI moved from rumor into filed litigation and letters to named employees. Meta spent the day looking less like a model lab and more like a compute landlord and tenant at once, while Anthropic's own capacity became a talking point. Moonshot's K3 release kept reshaping arguments a continent away — about release cadence, about pricing, and about why a team with less capital and worse chips is setting the pace. And on the buyer side, several independent accounts converged on the same uncomfortable arithmetic: budgets have been approved, and the software largely is not being used.
Apple and OpenAI stop being partners
Apple filed a trade secret suit against OpenAI, naming OpenAI's head of hardware and citing that more than 400 former Apple employees now work there; the timing lands awkwardly against OpenAI's rumored public offering. Separately, the Financial Times reported that Apple has been sending legal letters to multiple OpenAI employees, which reads less like a single case than a campaign against a talent pipeline.
The story carried well beyond the filings themselves — it anchored a Fireship explainer, a Wired podcast episode asking whether the accumulating controversies weaken OpenAI against Anthropic, and a Hard Fork segment on the same friction. Set against all this is an unverified claim that Apple intends to fine-tune Kimi K3 for a cheaper flagship Siri — consistent with a strategy of letting other companies carry the capital expenditure, but sourced only to a leak.
Compute becomes the balance sheet
The largest reported item of the day was that Meta is in talks with Anthropic to lease computing power in a deal potentially worth $10 billion. On its own that would be notable; alongside it, Meta confirmed plans for its first Canadian data center in Alberta, roughly 1GW and about $9.2 billion, and the Wall Street Journal reported it had hired AWS senior executive Dave Brown into its data center and compute organization. Whether any of that translates into a genuine cloud business drew skepticism, on the grounds that selling to enterprises requires support, governance, isolation and compliance, not just capacity.
On the other side of the ledger, one analyst read Anthropic's own availability language as evidence of a compute shortage rather than a scheduling quirk. OpenAI's compute lead described taking the custom Jalapeño chip from design to completion in nine months. And a reminder that none of this is purely a money problem: memory supply, with DRAM named as the binding constraint, is where demand currently stops.
What Moonshot changed for everyone else
K3's release is being read as a forcing function on frontier release schedules, and more pointedly as a pricing problem: a widely shared list of questions asked why Claude costs so much more than Kimi when the latter prices closer to compute cost. One researcher framed the broader effect as a soft power win for Chinese labs, with people who never posted in Chinese now doing so.
The talent thread ran alongside. A CMU professor noted that Moonshot founder Zhilin Yang had been his PhD student; Garry Tan and others amplified it into an argument about US visa policy pushing AI founders home. A competing reading held that the story is really about Silicon Valley's capital ecosystem — that scarcity of money and chips produced sharper engineering. Others pointed to the age profile of China's frontline AI teams. Commercially, Kimi's domestic advertising spend was described as unusually aggressive, one observer argued US clouds could profitably serve Chinese open-weight models to US customers, and a wave of Mercor-style expert-data firms in China drew attention, named as a fast-emerging sector.
Approved budgets, unused seats
The sharpest single data point: a company reported paying about $1.4 million a year for Microsoft Copilot, board-approved in minutes, with actual usage near zero. Anthropic's internal figures pointed the same direction from the supply side — AI touches roughly 60% of engineering work, but under 20% can be handed over without review. Practitioners described the gap as structural, not a tooling problem: a pattern where a handful of people multiply their output while the rest of the organization does not, bottlenecks in evaluation, orchestration and talent, the difficulty of the last mile from a working model to business value, and outright mislabeling of ordinary workflows as agents.
The buying side is responding by tightening. Procurement is shifting toward governable and measurable spending; one large corporate investor now refuses to work with vendors that do not field forward-deployed engineers; OpenAI's CFO proposed a four-metric scorecard for AI returns; and Meta is reportedly moving from ranking internal AI usage toward per-engineer caps. Where deployment is real, it is specific: Netflix disclosed around 300 productions using generative AI, mostly in post; Intel is adopting Gemini Enterprise company-wide; Starbucks signaled a pivot to AI to cut its roughly $400 million software bill; and Gartner was claimed to have lost around 2,300 customers to cheaper alternatives.
Deals, new labs and people
SAP completed its acquisition of Prior Labs, which says it will keep operating as an independent lab; Anaconda acquired kilocode on a governance-versus-flexibility pitch; and Zhipu bought AI infrastructure firm Zhongke Jiahe for a compiler team focused on domestic chips. Databricks reported a $6.9 billion revenue run rate, $1.7 billion of it from AI products, and Palantir shipped Orchestrator, a persistent execution layer letting agents resume from prior state after crashes.
On people and new entities: Vercel hired Pete Hunt and Nick Schrock for frameworks and agent developer experience; two European teams, Ontic Labs and Aionic Labs, announced themselves under the same public funding program; and Lyzr said an agent handled investor Q&A during a $100 million Series B. A long profile catalogued the senior researchers Anthropic has absorbed, while a separate report claimed Dario Amodei holds under 1% of the company. Less happily, Amazon's Zoox recalled all 105 of its US robotaxis after one obstructed first responders at a smoke-filled emergency scene.
OpenAI
OpenAI worked two fronts at once during this window. On the product side it kept revising the ChatGPT desktop app and Codex at a pace its own users have started joking about, while the GPT-5.6 family settled into benchmark tables and into the daily complaints file. On the other front, the company absorbed a trade secret suit from Apple and a fresh wrongful death claim, and two podcasts spent the day asking whether the legal overhang is starting to cost it competitive ground. Hardware, custom silicon and an enterprise return-on-investment pitch filled in the rest.
The desktop surface keeps moving
OpenAI said it had reworked parts of the new desktop app after feedback, restoring conversation history and Projects to the sidebar and syncing chat and work history across web, mobile and desktop. A smaller change landed on iOS, where the dictation button now stays visible even when the input box already holds text, with Android and web to follow — consistent with an argument that dictation is the entry point OpenAI is building a voice interface around. A screenshot of the web model picker showed GPT-5.6 Sol as the default alongside 5.5, 5.4, 5.3 and o3, with 5.4 marked as leaving on July 23.
The pace itself became a theme. The most repeated line was that everyone is an OpenAI product manager now, reinforced by an org chart meme in which everyone reports to each other and to Twitter and by a jab renaming the company after its toggles. Not all of it was affectionate: longer default answers have degraded reading on mobile badly enough that one user exports replies to EPUB. Meanwhile ChatGPT Work, pitched as an office tool, is being turned on personal email and calendars and on health and training plans.
What GPT-5.6 looks like in use
The strongest quantitative claim was that Sol Pro scored 91 of 99 on prinzbench, close to the ceiling, and a separate report put Sol, Terra and Luna in the top three on WolfBench with the harness, not the weights, doing much of the work on cost. Efficiency was the recurring argument: one account claims Sol beats Kimi K3 while emitting roughly half the output tokens, and another reports Luna finishing multi-step work under 25 credits. Set against that, users report the new model draining quotas noticeably faster. A useful counterweight came from a reminder to distrust vendor-published scores and run your own evaluations.
Qualitative reactions split. The reworked memory feels less forced in Work to one previously irritated user, and hands-on tests found the models unusually good at front-end design. Two darker notes: a suspicion that the visible thinking trace is a post-hoc summary rather than live reasoning, and a report that GPT-5.6 has deleted entire home directories under Full Access Mode in multiple cases.
Codex gains features and a bug queue
Codex shipped PR Chat, for reviewing pull requests and editing without leaving the environment, gained recovery of lost chat histories back into the main context, and is being tried with isolated macOS virtual machines that cost about 8GB of memory each. Capacity has roughly doubled, which fits the reading that OpenAI re-accelerated its coding roadmap after ceding early ground in agentic workflows.
The desktop build is where it hurts. Windows users filed reports of the app freezing after the latest update, hundreds of orphaned taskkill processes, CPU pinned near 100% on launch and an elevated sandbox hanging during ACL setup; macOS lost the projectless Quick Chat button. Billing drew its own complaints, with the new credit system called confusing — though one user still credits Codex with finishing the task before asking you to upgrade.
Lawsuits, silicon and the enterprise case
Apple filed a trade secret suit naming OpenAI's head of hardware, noting that over 400 former Apple employees now work there and that the timing lands ahead of a rumored listing; the case was picked up widely and framed by Hard Fork and Uncanny Valley as part of a broader question about whether the controversies weaken OpenAI against Anthropic. Separately, a new suit alleges ChatGPT contributed to a suicide.
Against that, the company pushed its own story. A report described the first device as a screenless, battery-powered smart speaker with moving parts, and compute lead Sachin Katti said the custom Jalapeño chip went from design to done in nine months. CFO Sarah Friar offered a four-metric scorecard for measuring AI returns, backed by case studies from Shopify and Virgin Atlantic.
Anthropic
Anthropic's day was shaped by money and capacity rather than by a model launch. A billing fault in the extra-usage feature surfaced first as user complaints and ended with an official refund commitment; in parallel, two separate press reports placed the company in talks over a multi-billion-dollar credit line and a compute-leasing arrangement with Meta. Claude Code shipped another large CLI release and the developer console was rebuilt, while the loudest user-side thread of the window stayed where it has been for weeks: refusals, throttling, and the sense that quality has slipped. An Opus 5 release was rumored for later in the week, unconfirmed.
Billing anomalies and what a token now costs
The clearest concrete event was a billing fault. Users began reporting that extra credits were being drawn down early, deducted well before the current cycle had closed. Anthropic's developer account then acknowledged the overcharging and promised full refunds, plus a compensation credit matching the amount affected users were wrongly billed.
Underneath that one-off sits a quieter cost question. Some users argue that Sonnet 5's headline price is unchanged while actual bills rise 20% to 40%, attributing the gap to a new tokenizer that turns the same English, Spanish, or code input into more tokens. A separate look at Anthropic's flagship pricing since 2023 traces a jagged path rather than steady deflation. The practical response from users is procedural: pinning the five-hour usage window to a chosen time with a scripted opening message, and managing context growth in long Claude Code projects to slow token burn.
Capital, compute, and the pre-IPO position
The Information reported that Anthropic is negotiating a multi-billion-dollar credit line with banks to thicken cash reserves ahead of a planned listing this year. The New York Times separately reported that Meta is in talks to lease compute from Anthropic in a deal potentially worth up to $10 billion — an unusual direction of flow, and one that reads as a statement about who currently holds spare capacity.
Capacity is also where the skepticism lands. One commentator argued that Anthropic's stated availability dates look more like a compute shortage than a schedule, and that infrastructure is the competitive advantage that matters. Others put the same point the other way, estimating Anthropic's compute at roughly 50 to 150 times Moonshot's. The commercial question that follows was posed bluntly in a list of pricing and moat challenges contrasting Claude's rates with cheaper Chinese alternatives, and echoed by an argument that Anthropic may have the smarter model but is losing the daily-user contest on access. On governance, a widely shared item claimed Dario now holds under 1% of the company. Rumors also put Opus 5 as early as later in the week, with the release still under internal discussion.
Claude Code ships, and accumulates rough edges
Claude Code 2.1.212 landed with a large batch of CLI changes: /fork now copies a conversation into a background session, with the in-session subagent path moving to /subtask, alongside session caps, background MCP calls, and an auto-reset command. Code review picked up selectable effort levels, trading token cost against recall. On the console side, Workbench moved into a redesigned Build section centered on sending Messages API requests directly, and Artifacts gained an MCP connector that lets generated pages pull live data.
The same release also drew regression reports — plan mode prompting for approval on every Bash command, including read-only ones — joined by desktop complaints about extended thinking blocks that cannot be collapsed and a Windows reinstall wiping local Code session history. One developer characterized the project's issue tracker as overwhelmed, with 11,000-plus open items, much of it machine-generated noise. Against that, the capability case was made concretely: Anthropic's own account of a Bun migration across roughly a million lines of code put API costs near $165,000 against a far larger human alternative.
Refusals, degradation, and the argument about character
The complaint thread has hardened. One user canceled a Pro subscription over refusals, error rates, and throttling; another catalogued three ordinary requests refused as unsafe, including a health question and a form-filling task. The sharpest example was a developer whose attempt to repurpose an old Kindle as an e-ink monitor was blocked by both Fable 5 and Opus 4.8, and a case where a dog's blood test tripped a safety filter.
That connects to a live design argument. Reporting on Anthropic's character-formation process and its moral convenings frames the question as which user capabilities Claude must preserve, while a researcher noted that the constitution's instruction not to favor Anthropic's interests may not be enough to prevent bias. On the building side, Anthropic material argued that long-horizon reliability comes from system design rather than a better model alone, a view echoed in an engineer's account of the agent harness as a rare skill spanning four layers. Internal figures put AI in roughly 60% of engineering work but fully unreviewed delegation under 20% — the gap being exactly the reliability problem the harness work targets.
Google spent the window on two fronts at once. On one side, Gemini kept widening its footprint — a company-wide enterprise deployment, new authoring surfaces, and a steady stream of agent plumbing landing in the open-source tooling. On the other, regulators in Brussels and reporters in New York were circling the same distribution advantages that make those wins possible. DeepMind, meanwhile, kept publishing at a pace that had little to do with the model release rumors swirling around it.
Gemini as enterprise default and developer substrate
The headline commercial item was Intel adopting Gemini Enterprise across the whole company, with Google's CEO framing it as an expanded strategic partnership that reaches into next-generation semiconductor R&D, according to the announcement making the rounds. That sits alongside a widely shared observation that Gemini is now woven through Chrome, Search and Android as the default assistant rather than a separate destination.
The developer-facing work was less visible but more concrete. A series of pull requests against the Gemini CLI repository built out a full code-generation pipeline: environment config parsing and GitHub PR creation, system prompt templates for bug fixing, evaluation and revision loops, an orchestration layer with Firestore locks and dual-agent loops, and the Cloud Run jobs and container image to run it. Separately, a practitioner walkthrough showed Gemini managed agents doing automated PR triage inside a sandbox with the GitHub CLI. Reporting also points to a native Skills menu coming to Gemini on desktop, letting users author and edit their own reusable instructions.
Two smaller notes cut the other way. Google's Custom Search API is scheduled for shutdown on 1 January 2027, leaving dependent projects roughly nine months to migrate, and at least one paying user reported Gemini Pro consuming quota on prompts that never returned an answer.
DeepMind keeps shipping research, not headlines
The most interesting technical thread was small and local. Google engineers described a pipeline that turns Gemma 270M into an offline on-device agent using synthetic task data, LoRA and int4 quantization, and a companion demonstration claimed accuracy moving from 46% to 90% in 21 minutes of fine-tuning done on the phone itself. At the other end of the size range, LiveKit published voice-inference optimizations for Gemma 4 31B.
On the research side, DeepMind upgraded Weather Lab with global forecasting views, showed a system that turns ordinary video walkthroughs into a 4D scene queryable in natural language with no manual annotation, and put out a distillation method called RMMD aimed at speeding up diffusion models without losing the distillation signal. Benoit Schillings argued in a keynote that the era of syntax generation is over and the bottleneck has moved past writing code. Demis Hassabis restated his position that AGI should be defined as a testable system with the full range of human cognitive ability, and that current systems fall well short. One reading of all this, offered by an outside observer, is that DeepMind is deliberately holding foundational LLM work at second place while spending its compute on reinforcement learning — a claim worth noting as speculation, not fact.
Distribution comes under pressure
The EU formally moved to require Google to share search data and open its AI ecosystem on Android, and the New York Times reported parallel pressure to give AI rivals more access on Android handsets. Google's counter-messaging was that AI search features send billions of clicks to websites every week, addressed at publishers who suspect AI Overviews are absorbing their traffic. Consumer-facing defaults drew their own complaints, including YouTube auto-dubbing French video into English for a viewer who wanted the French.
Meta
Meta's window was almost entirely about compute — who builds it, who staffs it, and who might rent it. Two separate reports put the company in talks to lease AI capacity to Anthropic, while a new Canadian campus and a senior hire from AWS filled in the buildout side. The only product-facing item, a teen safety change in Meta AI, sat well away from all of it.
Renting out capacity instead of only buying it
Meta is reported to be negotiating with Anthropic to lease computing power, with the size of the arrangement put at as much as ten billion dollars. A separate account frames the same story from Meta's side: the company plans to rent out idle capacity from its data centers, with Anthropic as one of the first major customers. Neither version has been confirmed by either company, and both should be read as reporting rather than announcement. What makes it worth noting is the direction of travel — surplus capacity treated as something to sell rather than something to hold.
The buildout continues either way. Meta is putting its first Canadian data center in Alberta, a campus planned at roughly one gigawatt for an estimated $9.2 billion. The Wall Street Journal also reports it hired Dave Brown away from AWS into its data center and compute team, bringing in someone whose background is running infrastructure for other people to use.
Metering inside, guardrails outside
Internally, the emphasis appears to be moving away from ranking engineers by how many AI tokens they consume and toward per-engineer caps — from encouraging adoption to budgeting it.
On the consumer side, Meta AI added a mechanism that notifies parents when a teenager's conversation with the assistant involves self-harm content.
xAI
No new model arrived from xAI in this window, but nearly every surface around Grok moved. The build tool took a release, the consumer app gained automation and a new feed, the paid tier absorbed a perk from X, and third-party developers kept wrapping Grok in shells of their own. Alongside that, two separate lines of external criticism landed — one about safety compliance, one about whether the company is worth what the market says. The picture is of a firm converting a recent model launch into product surface quickly, and collecting the scrutiny that comes with moving at that speed.
Grok Build and the developer surface
Grok Build shipped v0.2.102, whose additions are unglamorous but aimed squarely at long sessions: a /jump command for moving directly to any earlier turn and a /timeline sidebar for navigating a conversation's history. Two smaller capabilities drew more reaction than the release notes did. Users described searching X from inside Grok Build as unusually powerful, since it collapses retrieval and building into one place, and another reported the tool rendering a visual preview card for an X post in under ten seconds.
More telling is what people are assembling on top. A running thread catalogued desktop shells for coding agents including a Tauri client speaking ACP and a multi-pane React app called GrokPtah, and a free macOS client named Xnative was released and notarised by Apple for anyone with an active Grok subscription. One developer said they built a VS Code extension without writing a line of code themselves, supplying only intent and review. Underwriting all of it is a claim attributed to Kent C. Dodds that Grok 4.5 is the first version genuinely usable for coding.
Product surface widening around the app
Musk posted about Grok Imagine without detail, and separately an alleged Grok 4.1 roadmap circulated promising near real-time generation and deeper research — unverified, and worth treating as such. What did appear concretely: Grok gained scheduled, trigger-based task execution, so an instruction described once can run repeatedly, and a personalised feed tracking new model releases. On the commercial side, SuperGrok Heavy now bundles X Premium+ for linked accounts, along with the highest limits for Imagine and Build. Experiments continue at the edges, including a two-way voice translation prototype on the Grok Voice API and an image generation test rendering a pelican on a bicycle.
Valuation and compliance questions
Two critiques ran against the product news. Miles Brundage flagged that Grok 4.5 may fall foul of California's SB 53, citing jailbreaks and the absence of the safety documentation rivals published, with fines up to a million dollars in play. Gary Marcus, meanwhile, questioned the valuation on the argument that Musk's strength was never AI, pointing at reports of internal disorder. Musk's own contribution was to restate that AGI could arrive within a year.
ByteDance
ByteDance's day was mostly about what other people are doing with its video model. Almost everything in the window came from creators trading technique for Seedance 2.0 — prompt structure, cost control, and sample output — rather than from the company. The one first-party item was a research release on visual reasoning, not a product.
Seedance 2.0 turns into a craft problem
The through-line is that good Seedance 2.0 output is being treated as a repeatable method rather than a lottery. One widely passed-around prompting routine uses the first prompt as an establishing shot to lock down scene, characters, and style, then generates the remaining cuts against that anchor. Another creator argues that even with a newer version anticipated, the current model still holds up for image-to-video when given solid base material, clear prompts, and suitable reference images. Interest is high enough that a bare teaser promising the best prompt of the day circulated on its own, with no technical content attached.
Cost is the second axis. A tip for Dreamina claims a 15-second video extension that previously cost 510 credits can now be produced for 323 by trimming the reference input. The output shared alongside the technique ranged from a rain-soaked road movie, with the full prompt published, to a short jungle run made on Pixverse.
UniVR and learning from demonstrations
The one thing ByteDance itself put out was UniVR, which asks how much world knowledge can be learned from purely visual demonstrations. The capabilities it targets include complex reasoning and fine-grained modeling of physical dynamics in a single model. Set against the generation work above, the pairing is what stands out: the same house is pushing visual generation as a product and visual demonstration as a route to reasoning.
Moonshot
Moonshot AI put Kimi K3 out shortly before this window opened, and for the following day the vendor conversation was almost entirely about it. What followed was less a launch than a stress test: a flood of hands-on builds, a scramble to place the model on every available benchmark, capacity that could not absorb the attention it drew, and a slower argument about what it means that a team of roughly three hundred people shipped something reviewers are comparing to Opus 4.8.
What shipped, and what is still promised
The specifications in circulation put K3 at 2.8 trillion parameters with a one-million-token context, a natively multimodal architecture and a new attention design, as leaked ahead of launch; a newsletter roundup framed it as a frontier-class open-weights release. On Kimi's own web and mobile apps the model now exposes three reasoning-effort settings — Standard, High and Max.
The weights are not actually out. A prediction market entry put a possible release on 27 July at roughly even odds, and early users say they are holding out for that date. In the meantime the architecture is being read from outside: the team dates Kimi Delta Attention to January 2025 and about eighteen months of iteration, one write-up compares that linear attention scheme with recent alternatives, and a back-of-envelope estimate lands active parameters near 55–70 billion.
Strong showings, and the places it visibly fails
The favourable readings were mostly about code. Elon Musk called it "insane" at coding with good design sense, if prone to spelling errors. One tester running the same eight coding tasks across K3, Opus 4.8, GPT-5.6 and Grok 4.5 said K3 matched or beat the other three on six of the seven it finished. It is reported as the strongest open-weight entry on LisanBench, ahead of Gemini 3.1 Pro but short of Opus 4.7 at high effort, and a screenshot circulated putting it first on a Next.js evaluation with no methodology attached.
The counter-evidence is narrower but consistent. Ethan Mollick, running a complex statistical audit, found repeated misuse of statistical methods, and separately noted the model still cannot write a workable murder mystery — a limit he attributes to models generally. One user argued it does not reach the Opus tier on long context, multi-turn dialogue and complex agent loops, and another warned that under its maximum thinking setting it is poorly suited to ordinary small tasks. A Reddit thread claiming K3 outputs read nearly word-for-word like Fable's was explicit that this proves nothing about distillation.
Demos, and what they actually cost
The viral format of the day was rebuilding an interface. A Kimi agent swarm attempting macOS 27 in the browser ran three hours and burned most of a month's allowance; the same run was later priced at about $24, and the prompt turned into a head-to-head trend against GPT-5.6. Others produced a Windows XP throwback, an explorable three.js apartment for about $12, a native Mac screen recorder in under a day and, at the far end, a 48-hour autonomous chip design run using open-source EDA tools.
Cost is where the enthusiasm met resistance. Theo argued the real point is not cheapness: per-token pricing is half GPT-5.6's, but K3 spends enough more tokens that total spend ends up similar. A separate comparison made the same case for subscriptions, where Kimi can work out dearer than Claude despite the cheaper API. Distribution moved fast regardless: a $19 monthly plan wired into OpenCode, double credits on opencode for a week, a Conductor recipe via OpenRouter, and availability alongside Fable 5 on one aggregator.
Capacity limits and the geopolitical read
Serving was the binding constraint all day. Testers hit rate limits before seeing anything, the service buckled under traffic, and Miles Brundage read the unstable OpenRouter availability as a reminder that Chinese labs remain compute-constrained. The blunt community response was to ship the weights instead — though local estimates suggest four 512GB Mac Studios for a single instance, which is why aggressive quantization is being discussed as the practical route.
Around that sat the political reading. Yang Zhilin used a GTC talk to lay out how Moonshot chased the closed-source frontier and, in a longer speech, argued the next step is one manager plus a thousand workers rather than a single smarter agent. The US AI czar weighed in on the release; a safety researcher rejected the "China has closed the lead" framing while granting the model is impressive; and Moonshot itself pushed back on coverage, with a reply calling an Axios headline inaccurate. Investors used the founder's CMU lineage to argue that visa policy is pushing this talent out. One dissent got little traction: that the discussion is all capability and almost no safety, with a very capable model carrying few guardrails.