Solo dev hits Together AI rate limits running parallel agents, seeks generous API credits

Correct_Positive_108 · reddit · 2026-10-07

A solo developer building an agentic repository-indexing and benchmark-generation tool says Together AI's RPM/TPM limits became a blocker once they moved from toy scripts to running multiple parallel agents across Llama 3.3 70B and Qwen 2.5. The models themselves work fine—concurrency is the issue—and enterprise-tier pricing is out of reach without revenue, so they're asking for developer-friendly inference providers with enough API credits to experiment.

Original post →

More from Infra

Infra channel →