Automatic idle shutdown for dedicated deployments
UpdateUnverifiedAdded Sep 24, 2026
Deployments can now stop themselves when they go unused. Set an inactivity timeout with --inactive-timeout (the inactiveTimeout field in the management API), and if the deployment serves no inference requests for that many minutes, it scales to zero replicas, releasing its hardware and stopping billing. See Automatic idle shutdown .
Topics: Pricing, Developer tools
More Together AI platform releases
Every Together AI platform release| Date | Release | Type |
|---|---|---|
| Sep 15 | Rollouts for dedicated model inferenceunverified Beta | Beta |
| Sep 15 | Together CLI v2.34.0unverified Beta | Beta |
| Sep 15 | Model deprecationsunverified Deprecation | Deprecation |
| Sep 14 | New serverless modelsunverified Update | Update |
| Sep 11 | New serverless modelsunverified Update | Update |
| Sep 10 | Preemptible compute for GPU clustersunverified Preview | Preview |
| Sep 10 | What's new:unverified Update | Update |
| Sep 10 | Together CLI v2.33.2unverified Beta | Beta |
Also shipped on Sep 16, 2026
Together AI in September 2026Sources: each vendor's own release notes, changelogs and GitHub releases, read daily to monthly by how often it posts. Logos via logo.dev; trademarks belong to their owners.