Local Multimodal Model Reads Medical Reports
MaziyarPanahi · x · 2026-07-17
A demo of a locally running multimodal model:
- The model is named Inkling, uses open weights, and supports text, image, and audio.
- The author ran it on a Mac Studio using a quantized GGUF version and the llama.cpp Metal runtime.
- In the demo, the model first extracts 23 values from a medical exam photo, then uses this information to reason, determining the most urgent issue to be hyperkalemia accompanied by kidney problems.
Related event: Multimodal Model Inkling Tested Locally for Medical Report Reading(3 posts)→
More from Multimodal
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22