Team With 2x A100 Grant Open-Sources Roadmap for Distilling Offline Edge Models

Lerok-Persea · reddit · 2026-09-13

A team granted 3 months on 2x A100 80GB SXM4 (160GB VRAM via NVLink, 2TB RAM) by hessian.AI shares its roadmap on r/LocalLLaMA for building models that run fully offline on consumer smartphones/tablets for zero-connectivity environments like construction sites and disaster zones.

Project 1: Voice-to-structured-data. Dictation to schema-validated JSON for construction supervisors and EMS responders: generate synthetic data with unquantized 70B teachers (Llama-3.1-70B / Qwen-2.5-70B), fine-tune Distil-Whisper on real noise (diesel engines, jackhammers, sirens), distill into 1.5B–3B models (Qwen-2.5 / Llama-3.2), quantize INT4/GGUF/AWQ for a sub-2GB mobile footprint.

Project 2: Blueprint wall detection. Automatic segmentation of walls and dimensions in 2D construction drawings, weighing direct VLM fine-tuning (Qwen2-VL / Florence-2 with bounding boxes/polygons) vs a hybrid SAM / YOLOv8-seg + small LLM pipeline.

They also crowdsource ideas: overlooked edge-AI niches (privacy auditing, industrial sensors, field inspections), novel 1B–3B distillation experiments, and tooling advice (vLLM offline batching + Axolotl / Unsloth).

Original post →

More from Infra

Infra channel →