Team With 2x A100 Grant Open-Sources Roadmap for Distilling Offline Edge Models
Lerok-Persea · reddit · 2026-09-13
A team granted 3 months on 2x A100 80GB SXM4 (160GB VRAM via NVLink, 2TB RAM) by hessian.AI shares its roadmap on r/LocalLLaMA for building models that run fully offline on consumer smartphones/tablets for zero-connectivity environments like construction sites and disaster zones.
Project 1: Voice-to-structured-data. Dictation to schema-validated JSON for construction supervisors and EMS responders: generate synthetic data with unquantized 70B teachers (Llama-3.1-70B / Qwen-2.5-70B), fine-tune Distil-Whisper on real noise (diesel engines, jackhammers, sirens), distill into 1.5B–3B models (Qwen-2.5 / Llama-3.2), quantize INT4/GGUF/AWQ for a sub-2GB mobile footprint.
Project 2: Blueprint wall detection. Automatic segmentation of walls and dimensions in 2D construction drawings, weighing direct VLM fine-tuning (Qwen2-VL / Florence-2 with bounding boxes/polygons) vs a hybrid SAM / YOLOv8-seg + small LLM pipeline.
They also crowdsource ideas: overlooked edge-AI niches (privacy auditing, industrial sensors, field inspections), novel 1B–3B distillation experiments, and tooling advice (vLLM offline batching + Axolotl / Unsloth).
More from Infra
- Skip Vercel: run your AI agent straight on a DigitalOcean droplet for instant deploys — Daniel_Farinax · 2026-09-13
- A silent CPU shortage is hitting cloud reservations for mid-size companies — Sethwinterroth · 2026-09-13
- DeepSeek V4.1 Flash Hits 40 tok/s on 8x A40 with Open-Source TensorSharp Engine — fuzhongkai · 2026-09-13
- Samsung's first 2nm Taylor fab fully booked by Tesla, Arm and Broadcom pre-launch — Beth_Kindig · 2026-09-13
- Reddit theory: the AI 'slowdown' talk masks a coming compute, power and datacenter crunch — maxpayne07 · 2026-09-13
- Prediction: hyperscalers will tighten AI capex in 2027 — pdamodaran · 2026-09-13