Why Enterprise Data Isn't Always Fit for Model Training
inductionheads · x · 2026-07-14
The post argues that a company's most valuable and sensitive data is often unsuitable for direct model training.
Key points include:
- Sensitive information typically cannot be freely fed into models due to strict access controls
- Building security boundaries directly into models or agents is highly challenging
- While custom model training will become common in many domains, training a bespoke model for every enterprise is far harder than it sounds
- A more viable approach is transforming enterprise knowledge into skills or artifacts that models can invoke in-context
- Model training is expensive and time-consuming, often requiring retraining after foundation model updates, with irreversible results
The discussion centers on the trade-offs between training data into models versus leveraging it via in-context calls.
Related event: Opinion: Most Enterprises Shouldn't Train Their Own Models Yet(3 posts)→
More from AGI Musings
- Frontier labs are concentrating AI safety expertise, and that may be distorting the debate — ohlennart · 2026-07-21
- Real AI progress in math often comes from counterexamples to old beliefs — LucaAmb · 2026-07-21
- AI is cutting costs faster than it is creating new revenue — kevinkern · 2026-07-21
- AIFEC may be the most credible way for humanity to steer its own future — gleech · 2026-07-21
- Personal AIs could make websites context for agents, not pages for humans — yacineMTB · 2026-07-21
- A screenshot revisits GPT-4’s napkin-to-code demo and imagines GPT-N building GPT-N+1 — genmon · 2026-07-21