Inkling Runs 2-Minute Consultation Audio Locally
MaziyarPanahi · x · 2026-07-16
A user reported further testing Thinking Machines' Inkling with excellent results. Follow-up details reveal they ran the model on a Mac Studio using a 2-minute doctor consultation audio sample.
Key takeaways from the post include:
- The model processes real audio directly, rather than relying on transcribed text
- While the initial complaint mentioned was a "knee issue," the model successfully identified and flagged the actual critical risk: heart failure
- The parameter scale is listed as 975B params
- It utilizes the 1-bit GGUF version
- It runs locally on llama.cpp, ensuring "nothing leaves the machine"
Overall, this is a hands-on showcase of local inference for massive models and native audio comprehension.
Related event: Testing 975B Inkling Model Locally for Accurate Audio Processing(4 posts)→
More from coding & agent
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11