CHANNEL
Infra
"Infra" is a topic channel on AGI Hunt, an AI news site updated around the clock in real time. Coverage: Compute, GPUs and chips, data centers, inference and serving stacks, training costs, supply chain and compute economics.
Daily roundup: the latest AI News Daily — the past 24 hours across the whole site, per channel and per company · browse the archive
- Yacine Warns AI Platforms Can See Your Work, Calls for Sovereign AI — yacineMTB · 2026-09-08(3 related)
- Crusoe raises $3B at $30B valuation as it builds Stargate, delivered 200MW in 11 months — 快鲤鱼 · 2026-09-08
- CPU-Only LLM Tests: 35B MoE at Q2 Beats a 2B Model Despite Half the Speed — ML-Future · 2026-09-08
- HydraDB: a Rust graph database that lives entirely in S3, with nothing on disk — thisdudelikesAI · 2026-09-08
- Running Qwen 27B and Gemma 31B locally on one RTX 4090: quantization and context tradeoffs — MooseEfficient2151 · 2026-09-08
- Could distributed iPhones form an inference network? A 2AM open question — gajesh · 2026-09-08
- Candidate depth 10→500 lifts score by just 0.01: Qdrant on diagnosing before tuning vector search — qdrant_engine · 2026-09-08
- Custom llama.cpp build pushes 7900XTX to 1600tk/s prefill on Qwen 27B Q8 — nasone32 · 2026-09-08
- Discrete Diffusion Parallel Sampling Delivers Lossless LLM Speedups Without Draft Models — IFM · 2026-09-08
- Coding Harness Matters More Than Model: Token Use Varies 9.2x — zainhas · 2026-09-08(2 related)
- JapanFold launches: 8 open biology AI models served free for Japanese researchers — DavidBennett__ · 2026-09-08
- Running a local MLX music model via Codex pushes M4 MacBook Air to its limits — Dimillian · 2026-09-08
- Qdrant releases 10B-vector retrieval dataset with exact top-1000 ground truth for 100K queries — qdrant_engine · 2026-09-08
- Debunking the viral 'data center infrasound harm' videos: citations don't hold up — AndyMasley · 2026-09-08
- AI cluster builders bypass the grid: gas turbines, fuel cells and SMRs come onsite — AccBalanced · 2026-09-08
- Jensen Huang Calls GPUs Income-Producing Assets as Aging H100 Rents Jump 22% — AccBalanced · 2026-09-08(2 related)
- Inside the VMs Powering Mobile AI Agents: Instinct, Claude Code — RohanAdwankar · 2026-09-08
- FT: CXMT and YMTC stockpiled enough ASML DUV tools for three years of expansion — basedjensen · 2026-09-08
- Arm unveils Mali G2-Ultra NX GPU: desktop-class mobile gaming with AI-native graphics — Re-Tails · 2026-09-08
- Nvidia's Quarterly Profit of $59.7B Tops 12 Corporate Giants Combined — FinanceYF5 · 2026-09-08(2 related)
- Consolidating four small models into one inference server: a doc-QA agent's ops tradeoffs — Sad-Razzmatazz-7657 · 2026-09-08
- Huawei Ascend takes PyTorch China stage: from hardware adaptation to joint standards research — PyTorch · 2026-09-08
- AUR llama.cpp-cuda removal caused 10x slowdown; manual rebuild restored 1800 t/s prefill — MrHall · 2026-09-08
- CPO Laser Debate: MOPA Architecture vs Lumentum's Approach — vikramskr · 2026-09-08(2 related)
- DRAM revenue surges 385% YoY; Apple's foldable iPhone lands next week — 创业邦 · 2026-09-08
- Longsys lists in Hong Kong, raising $910M as AI-driven memory boom fuels record profits — 创业邦 · 2026-09-08
- Report: Anthropic Walks Away From $6B Decart Acquisition After Due Diligence — ns123abc · 2026-09-08(6 related)
- VDN-H3 video generation running on 24GB VRAM: 30-second clips on an RTX 3090 Ti — NoMouse9610 · 2026-09-08
- NeoCloud Operators Face Harsh Contracts as Rack Failure Could Cost Six Months of Rent — rohanpaul_ai · 2026-09-08(2 related)
- Radix-select top-k fix boosts Qwen Flash-Next decode 9-12% at ~119k context on 2x3090 — Extension-Bid-639 · 2026-09-08
- All Roads Lead to Neocloud: Seven Types of Compute Companies Converge — FinanceYF5 · 2026-09-08(2 related)
- Cambricon and Alibaba Cloud Join PyTorch Foundation as Platinum Members — PyTorch · 2026-09-08(4 related)
- AI cancer cures slowed by chip shortage, says Arm boss — dabinat · 2026-09-08
- Arm's CSS for Mobile 2 Skips the NPU, Bets On-Device AI on CPU and GPU Engines — ryanshrout · 2026-09-08
- fal H3 Max Delivers Real-Time 1080p Video Generation — isidentical · 2026-09-08(4 related)
- CPU Shortage Reaches Software Teams as AI Chip Supply Constraints Spread Beyond AI — brada · 2026-09-08
- Tech Giants Eye Argentine Patagonia for Massive Data Centers — Polymarket · 2026-09-08(3 related)
- TimesFM 3.0 merges native MLX backend: 641 series/sec on M4 Max, no PyTorch needed — rachittshah · 2026-09-08
- Data center buildout is driving freight demand that the Cass index misses — kernelangus420 · 2026-09-08
- Strix Halo users ditch official llama.cpp: optimized forks hit ~60 t/s decode vs ~20 t/s — feelspeaceman · 2026-09-08
- AMD Unveils Desktop AI Workstation With Triple DGX Memory — shashib · 2026-09-08(2 related)
- DavidSHolz: AI services should let users pay extra to spin up CPU/GPU VMs — DavidSHolz · 2026-09-08
- What broke running an agent fleet 24/7 — it wasn't prompt quality — ParrotIntegrated · 2026-09-08
- vLLM adds Hybrid HiSparse offloading to keep GLM 5.3 decoding past GPU memory — vLLM Blog · 2026-09-08
- Dual RTX 3060 vs single RTX 3090: choosing a $500 vs $1000 build for a home local LLM server — Bakkario · 2026-09-08
- Huawei 950DT reportedly priced at ~$16K per chip in DeepSeek order, below expectations — teortaxesTex · 2026-09-08
- kernel.org burns 14 CPU cores just rendering pages for abusive AI crawlers — Simon Willison · 2026-09-08
- Report: Intel to raise PC CPU prices 10%, may cut another 5-10% of workforce — zephyr_z9 · 2026-09-08
- TPU shipment forecasts for 2027 diverge: Morgan Stanley says 7.4M, UBS 10.4M units — Beth_Kindig · 2026-09-08
- WSL Manager 2.0 ships a built-in MCP server letting agents create and run Linux distros — bostrot · 2026-09-08