CHANNEL
Infra
"Infra" is a topic channel on AGI Hunt, an AI news site updated around the clock in real time. Coverage: Compute, GPUs and chips, data centers, inference and serving stacks, training costs, supply chain and compute economics.
Daily roundup: the latest AI News Daily — the past 24 hours across the whole site, per channel and per company · browse the archive
- NVIDIA Claims Nemotron 3 Ultra Hits 97.1% in Chip Design — NVIDIAAI · 2026-07-27(2 related)
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- MiniBot 2.40 adds xAI, HF Studio and vLLM support with inline media tools — Creative-Type9411 · 2026-07-27
- Apple smart glasses, Nvidia-SK AI data center deal, and Ctrip’s RMB 5.179 billion fine headline a tech roundup — APPSO · 2026-07-27
- DeepSeek funding rumor, EU AI transparency rules and OpenAI agent incident make a packed AI news roundup — 创业邦 · 2026-07-27
- QuixiCore argues native quantized kernels beat dequant-then-generic execution — QuixiAI · 2026-07-27
- Nvidia Reportedly in Talks to Back OpenAI's $250B Ohio Data Center — Wonderful_Buffalo_32 · 2026-07-27(8 related)
- Google’s JAXBench benchmark lifts TPU kernel optimization with 50 real workloads — omarsar0 · 2026-07-27
- TickerTrends says Anthropic’s ARR now tops OpenAI’s by $32.8B — wen_ragnarok · 2026-07-27
- Intel says 18A production has started and 14A is moving toward the next node — BenBajarin · 2026-07-27
- With an RTX 3090 Ti, should you run models in fp16, fp8, or int8? — throwaway0204055 · 2026-07-27
- Modal GPU rendering runs Blender scenes for under a cent, with L40S fastest in test — charles_irl · 2026-07-27
- Intel says server CPU ASPs rose 48% in its latest 10-Q — BenBajarin · 2026-07-27
- World Model Optimizer launches a router that cuts agent inference cost by 40%+ — SilenN · 2026-07-27
- Alphabet reportedly backs leases for 2.4GW of AI capacity across 10 projects — BenBajarin · 2026-07-27
- Bristol Myers plans a pharma AI supercomputer with Nvidia — owl_posting · 2026-07-27
- Reddit asks whether the AI spending bubble is about to pop—and cheap RAM might be next — RuiRdA · 2026-07-27
- AI-Trader adds an MCP server so LLMs can run trading backtests — tom_doerr · 2026-07-27
- New Q8_CR GGUF format keeps Krea 2 diffusion speed near INT8 while shrinking VRAM pressure — molbal · 2026-07-27
- Nvidia signs $1.5B multi-year Amkor deal to expand US chip packaging capacity — Beth_Kindig · 2026-07-27
- PyTorch DDP misses a tiny-parameter NVIDIA GPU optimization out of the box — gordic_aleksa · 2026-07-27
- Fully Local AI Girlfriend Voice Demo Runs on 15GB VRAM — max_paperclips · 2026-07-27(2 related)
- OpenAI reportedly plans to spend over $30B on a 3.2GW Georgia data center — Beth_Kindig · 2026-07-27
- Two RTX 5080s failed to beat one RTX 5090 in ComfyUI single-render tests — Geekdomo · 2026-07-27
- Open-source profiler tracks every STT, LLM, and TTS call in self-hosted voice agents — mahimairaja · 2026-07-27
- SemiAnalysis: Memory and Speed Trump GPU Power in AI Inference — rwang07 · 2026-07-27(2 related)
- llama.cpp merges GLM-5.2-Vision support for local multimodal inference — QuixiAI · 2026-07-27
- Cheap storage makes SCD Type 2 look obsolete, says a Meta-style data engineer — Zachly · 2026-07-27
- After Copilot went unlimited, one Reddit user is deciding whether to sell a two-GPU AI rig — tweetibird · 2026-07-27
- Open models rose from 10% to 30% of tokens in a year, Together says cost is the reason — togethercompute · 2026-07-27
- AI storage meme says the “marketing answer” is 5.3, but the real benchmark is still undefined — JoshuaJBouw · 2026-07-27
- NVIDIA says AdamW hits a scale ceiling as SOAP and Muon beat it on trillion-token runs — omarsar0 · 2026-07-27
- Local LLM builders ask whether RTX Ada workstation cards are worth tracking — egudegi · 2026-07-27
- Guide breaks down how to cut Microsoft Fabric capacity costs on Azure — adnan_hashmi · 2026-07-27
- Introductory guide explains the `sempy.fabric` package for Microsoft Fabric — adnan_hashmi · 2026-07-27
- Voice AI’s real call cost is more than minutes: STT, TTS, SIP and retries — decant338 · 2026-07-27
- Micron and Meta paper says Spark can slow down 38× when shuffle spills to SSD — dr_alphalyrae · 2026-07-27
- Inference businesses may be charging 7× to 15× more than renting a GPU — JoshPurtell · 2026-07-27
- CUDA benchmark shows 64M-number sum runs 20× faster on GPU than CPU — ctjlewis · 2026-07-27
- AI buildout pushes data-center financing into more creative capital structures — GaryMarcus · 2026-07-27
- Users ask whether RAID 0 NVMe setups improve local large-model runs on Pulsar — Wyldkard79 · 2026-07-27
- GLM-5.2 inference on RTX 5090s jumps from 30 tok/s to 80–110 tok/s — markjeffrey · 2026-07-27
- Hermes Agent says progressive tool disclosure scales MCP tools with near-zero accuracy loss — Teknium · 2026-07-27
- Laguna tests 2.75 and 3.25 bpw quantization with NVFP4 experts and FP8 KV cache — QuixiAI · 2026-07-27
- HART OS is an open-source AI operating system aimed at datacenter-free frontier AI — hevolveai · 2026-07-27
- Chutes says it trained a 20B model for under $10 an hour using rented GPUs across two continents — markjeffrey · 2026-07-27
- llama.cpp Adds Minimax-M3 Vision and MSA Support — Time_Reaper · 2026-07-27(2 related)
- AST-grep rewrites Tree-sitter in Rust and claims a 30% speedup — herrington_d · 2026-07-27
- Windows 11 ComfyUI user gets Sage Attention and FlashAttention 2 running on RTX 5090 FE — Left_of_Laniakea · 2026-07-27
- Baseten Pushes GLM-5.2 to 280 tok/s — iamrobotbear · 2026-07-27(3 related)