CHANNEL
Infra
"Infra" is a topic channel on AGI Hunt, an AI news site updated around the clock in real time. Coverage: Compute, GPUs and chips, data centers, inference and serving stacks, training costs, supply chain and compute economics.
Daily roundup: the latest AI News Daily — the past 24 hours across the whole site, per channel and per company · browse the archive
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek Releases V4.1-Flash with 1M Token Context — matlabulous · 2026-09-11(2 related)
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next Hexagon NPU to run 30B MoE models on-device — lee_stott · 2026-09-11(2 related)
- Stanford Paper: Hybrid Local-Cloud Routing Cuts 60-80% of AI Costs — rohanpaul_ai · 2026-09-11(2 related)
- Rented GPU Bills: Host CPU and Script Defaults Made Costs 31x Higher — Worldly_North_7213 · 2026-09-11
- Routing NVIDIA PAIR to llama.cpp on an AMD ROCm node (2×R9700): full notes — Don_Reuter · 2026-09-11
- Running Qwen3.8 locally on a 128GB laptop for agentic coding: thinking tokens, not tok/s, set the wall clock — deepu105 · 2026-09-11
- SpaceX Plans to Make Scarce Turbine Parts as AI Data Centers Strain Power Supply — rohanpaul_ai · 2026-09-11(2 related)
- AGI as task time horizon vs meetings — and why fabs should train their own models — jwt0625 · 2026-09-11
- Free client-side calculator compares LLM token economics across DeepSeek, Claude, o3-mini — nikola_mr64990 · 2026-09-11
- We Replaced mmap with io_uring in Our Rust Query Engine — It Got Slower — rzk · 2026-09-11
- 10 resources on what happens after training: KV-cache, quantization, serving — techNmak · 2026-09-11
- Full-precision DeepSeek 4.1 Flash hits 300+ TPS on 4 RTX Pros with custom vLLM fork — TheZachMueller · 2026-09-11
- Inside a Trillion-Token/Day Factory: Mooncake Turns KVCache into a Cluster-Shared Pool — 量子位 · 2026-09-11
- Web Demo Approximates V4.1 Flash-Style Fast KV Prefill on Qwen3 — T_rex2700 · 2026-09-11
- OpenAI CFO: Compute Bought a Year Ago Now Worth 3-5x — rwang07 · 2026-09-11(3 related)
- iFlytek's Spark X2.5 trained on 10,000 domestic Ascend 910B GPUs with 97% uptime — 机器之心 · 2026-09-11
- REVA Mines LLM Attention into Reusable Evidence Views, Cutting RAG Compression Overhead up to 15.6x — _reachsumit · 2026-09-11
- Edge0-35B-A3B preview MoE model for edge inference trends on Hugging Face — Edge0 · 2026-09-11
- Reflect Orbital Plans to Sell Sunlight via Satellite Mirrors — kyliebytes · 2026-09-11(3 related)
- Data centers are for startups, not frontier labs: more compute is the anti-monopoly move — arthurcolle · 2026-09-11
- DeepSeek Compresses KV Cache 54x in Nine Months to Sub-KB Per Token — max_paperclips · 2026-09-11(6 related)
- Anatomy of Jensen Huang's 'AGI Is Here' Tweet: Zero-Cost Signaling and Abilene's 357-Job Data Center Deal — 创业邦 · 2026-09-11
- China's CXMT plans four new DRAM fabs by H2 2028, closing in on Samsung and SK hynix — zephyr_z9 · 2026-09-11
- Dev questions whether OpenAI's prompt_cache_key design wastes massive compute — YouJiacheng · 2026-09-11
- SF Compute signs $245M in take-or-pay contracts for NVIDIA Blackwell B300 capacity — mattshumer_ · 2026-09-11
- SpaceX CFO: vertical integration is core, Starship paves way for orbital compute — elonmusk · 2026-09-11
- Bezos: Power Supply Chain Bottleneck Forces AI Labs to Slow Development Pace — beffjezos · 2026-09-11
- vLLM v0.29.0 Cuts Blackwell Latency 33.6%, Model Runner V2 Default — vllm_project · 2026-09-11(4 related)
- One cheeseburger emits as much CO2 as 63,000 Gemini text prompts, math shows — recallingmemories · 2026-09-11
- Google signs deal to buy half the electricity of a nuclear power plant — lukaspetersson · 2026-09-11
- spcx reportedly signed another mega compute deal a week ago, $13B ARR per CFO — rwang07 · 2026-09-11
- PyTorch Lightning checkpointing runs up to 95% faster on Google Cloud — LightningAI · 2026-09-11
- LuxTTS: open-source voice cloning at 48kHz, 150x realtime, under 1GB VRAM — tom_doerr · 2026-09-11
- IFM Open-Sources K2 Horizon: Six Models from 0.9B to 375B with Parallel Decoding and Benchmark Cheating Audit — rohanpaul_ai · 2026-09-11(7 related)
- Nearly 10% of exposed LiteLLM gateways accept default admin key 'sk-1234' — Thionne_WTZ · 2026-09-11
- DOJ probes Nvidia's $17B license-and-hire absorption of Grok rival Groq — Servola-Journal · 2026-09-11
- Nari Labs launches 50ms TTS endpoint, 10x cheaper than ElevenLabs on Qwen3-TTS — iamaliveix · 2026-09-11
- Amkor raises Arizona advanced packaging investment to $12B from $7B on surging demand — zephyr_z9 · 2026-09-11
- Qwen 3.8 Flash Next Optimization Challenge: MLX and CUDA Both Gain Over 55% — gajesh · 2026-09-11(2 related)
- 26 LLM Routers Caught Injecting Malicious Tool Calls and Stealing Credentials, One Client Lost $500k — RexDouglass · 2026-09-11
- DIY-friendly KiCad footprints for AI MELF resistors, milled at home — debreuil · 2026-09-11
- System76 launches Thelio Mira AI Linux workstation with 192 GB GPU memory — jonifico · 2026-09-11
- Inference providers barely break even: $10K revenue yields just $200 profit — metalvendetta · 2026-09-11
- Colocated async RL gains steam as observers speculate k3 uses it too — stochasticchasm · 2026-09-11
- Peter Diamandis: The AI race is becoming the biggest construction project of our generation — PeterDiamandis · 2026-09-11
- k3 report details: 50M sandboxes and reasoning-effort control — stochasticchasm · 2026-09-11(2 related)