Crash-testing ChatGPT plugin discovery with 3,000+ queries: small catalog, rarely invoked
Alpic-ai · reddit · 2026-10-01
Alpic ran over 3,000 queries against OpenAI's newly announced plugin search to stress-test it.
Findings: the catalog is small and skewed toward business software, and models rarely consult it unless the user explicitly asks. Their article details the methodology and implications.
More from coding & agent
- Translating an entire book with DeepSeek: pennies and under an hour, decent quality — teortaxesTex · 2026-10-02
- OmniSeek turns Omni-LLMs into agents that actively seek audio-visual evidence — Haibo Wang · 2026-10-02
- Microsoft's ActiveSaddler Uses Automated Curriculum Learning to Boost Agent Harnesses by 7.5 Points — microsoft · 2026-10-02
- Alibaba's PoS Maintains Explicit Belief States to Fix Long-Horizon Agent 'Belief Trapping' — alibabagroup · 2026-10-02
- IntentFlux Benchmarks 'Intent Drift' in LLM Agents: Scores Fall from 0.476 to 0.384 as Users Change Their Minds — Yanjie Zhang · 2026-10-02
- Founder says 36 hours with OpenAI dots may replace his months of monorepo agent setup — hugobowne · 2026-10-02