Flama 2.0: 8-year-old Python framework now packages LLMs into one file serving OpenAI, Anthropic and Ollama APIs
p3rdy · reddit · 2026-10-07
The author has developed Flama since 2018 with one goal: taking a trained model to a production API with minimal ceremony. It has served scikit-learn, PyTorch, TensorFlow and Transformers models, and version 2.0 makes LLMs first-class citizens.
Core design: everything is a .flm file. Every model, classic or generative, is packaged into a single .flm containing weights, auxiliary inference files (tokenizer, config) and metadata. The inference engine isn't baked in, so the same file runs anywhere; metadata sits ahead of the weights so flama model inspect reads it without loading, useful for CI gating.
One model, many protocols. Once added to the app, the model simultaneously answers on OpenAI (/llm/openai/v1/chat/completions), Anthropic (/llm/anthropic/v1/messages) and Ollama routes, with a built-in streaming chat UI. Anything that lets you set a base URL — harnesses, agent frameworks, SDKs — can talk to it. flama serve runs it straight from the CLI.
Three decoupled layers. A request flows through a dialect (OpenAI/Anthropic/Ollama/native) → a canonical transport → a backend (vLLM on Linux CUDA, MLX on Apple Silicon). Since layers only meet at the canonical middle, engines and dialects are independent, which is why tool calls and token usage come out right in every protocol.
The framework also ships: type-hint-based dependency injection, schemas with Pydantic/Marshmallow plus generated OpenAPI docs, CRUD APIs from SQLAlchemy tables, DDD building blocks (repositories, units of work), JWT auth, turning any app into an MCP server, flama upgrade codemods, and a Rust core for routing, serialization, parsing, crypto and compression. Full architecture is covered in an arXiv paper.
More from coding & agent
- 30 个真实业务工作流清单:别把 Agent 测试搞复杂了 — VibeMarketer_ · 2026-10-08
- Hybrid agent pattern: cloud Gemini plans, local Gemma swarm runs 97% of tokens offline — clmt · 2026-10-08
- Agent compaction drift: limits cascade into agents policing cgroups instead of doing the work — generativist · 2026-10-08
- Long-running agents suffer 'constraint amplification': a subtle form of context rot — generativist · 2026-10-08
- Nautilo Ships Text+Vision Model Split, Preps Price/Security-Based Model Routing Gateway — Dan_Jeffries1 · 2026-10-08
- A fine-tuned 9B beats a 31B model: 600 labels, $0.12, 91% accuracy — julsimon · 2026-10-08