Zilliz CTO: agents make the enterprise data layer impossible to ignore
No_Engineer_1224 · reddit · 2026-09-24
James Luan, CTO of Zilliz (the company behind open-source vector database Milvus), argues after Databricks Summit that the AI stack has evolved unevenly: model progress and compute demand are visible, but data work — duplicate records, inconsistent business meaning, stale pipelines, permissions scattered across systems — is easy to overlook.
Key points:
- Pre-training cost: preparing training data can consume days before an experiment starts; one wrong field or filter can invalidate everything downstream, and better compute doesn't remove the need to verify input trustworthiness.
- Agents amplify the problem: a capable model can act on an expired document or a record it shouldn't access; the operational question becomes whether context was valid at decision time and whether the system can trace where mistakes entered.
- Data infra grows more valuable: turning enterprise data into usable context requires quality, freshness, permission boundaries, and an affordable path into the agent's loop — ongoing operating responsibilities, not one-time cleanup after the first demo.
He also outlines Vector Lakebase, Zilliz's direction of bringing the vector serving layer closer to the data lake.
More from coding & agent
- QueueSim MCP Connector Lets AI Agents Run M/M/c Queue Simulations Across Four Scenarios — modelcontextprotocol · 2026-09-24
- predictfun-mcp Gives AI Agents Structured Access to $1.5B+ Prediction Market on BNB Chain — modelcontextprotocol · 2026-09-24
- Perf engineer praises memory-overcommit evals, suggests exposing stats to fleet scheduler — sloppenheimer · 2026-09-24
- Sandboxing AI agents with FnCall: network-only syscall limits and the MIG gotcha — sloppenheimer · 2026-09-24
- Merge launches Universal Context Layer, a shared company brain for every AI agent — shensi · 2026-09-24
- Lydia Hallie: Claude Design produced all video visuals in ~1.5 hours — lydiahallie · 2026-09-24