Solo dev hits Together AI rate limits running parallel agents, seeks generous API credits
Correct_Positive_108 · reddit · 2026-10-07
A solo developer building an agentic repository-indexing and benchmark-generation tool says Together AI's RPM/TPM limits became a blocker once they moved from toy scripts to running multiple parallel agents across Llama 3.3 70B and Qwen 2.5. The models themselves work fine—concurrency is the issue—and enterprise-tier pricing is out of reach without revenue, so they're asking for developer-friendly inference providers with enough API credits to experiment.
More from Infra
- The 2019 Mac Pro with 1.5TB RAM would be the ultimate local LLM machine today — Odd-Capital-847 · 2026-10-07
- Team claims sub-5-second full weight sync for 1T-parameter RL training — saurabh_shah2 · 2026-10-07
- NVIDIA's UNREAL: one model unifies corpus retrieval and long-context at 128K+ — nvidia · 2026-10-07
- SlimWise prunes MoE experts only at decode, boosting throughput up to 1.81x — Gunho Park · 2026-10-07
- Ora scanned 107,797 sites: average agent-readiness score just 41/100 — EdenEmarco177 · 2026-10-07
- Fine-tuned DFlash 2 Drafter Boosts Ternary Bonsai 2 27B by 2.2x on an L4 — naklitechie · 2026-10-07