Inference Doubled serverless rate limits for Small models We doubled the adaptive rate-limit ceilings for...
UpdateyesterdayUnverifiedAdded Oct 8, 2026
Inference Doubled serverless rate limits for Small models We doubled the adaptive rate-limit ceilings for Small serverless models (less than 600B total parameters). The Small tier now has these ceilings: Total Prompt TPM: 216M Uncached Prompt TPM: 54M Generated TPM: 2.16M Medium and Large model ceilings are unchanged.
More Fireworks AI platform releases
Every Fireworks AI platform release- Recommended migrations
Updateunverified
| Date | Release | Type |
|---|---|---|
| Oct 26 days ago | Update | Update |
| Oct 17 days ago | Update | Update |
| Oct 17 days ago | Update | Update |
| Oct 17 days ago | Update | Update |
| Oct 17 days ago | Update | Update |
| Sep 2810 days ago | Recommended migrations unverified Update | Update |
| Sep 182 weeks ago | Pricing | Pricing |
| Sep 182 weeks ago | Update | Update |
Also shipped on Oct 7, 2026
Fireworks AI in October 2026- SSH for Cloud Run services and instances is in Preview
Google CloudPreviewunverified
- Set credit usage limits and alerts for your workspace
LovablePricingunverified
- Faster inline text edits
LovablePreviewunverified
- Gemini 3.7 Flash is deprecated for your app's AI features
LovableDeprecationunverified
- The Cost Analysis tab of the Billing page now includes Serverless Inference Live Usage, which lists the...
DigitalOceanPricingunverified
| Date | Company | Release | Product | Type |
|---|---|---|---|---|
| Oct 7yesterday | Google Kubernetes Engine | GA | ||
| Oct 7yesterday | Cloud Run | Preview | ||
| Oct 7yesterday | Lovable | Pricing | ||
| Oct 7yesterday | Faster inline text edits unverified | Lovable | Preview | |
| Oct 7yesterday | Lovable | Deprecation | ||
| Oct 7yesterday | DigitalOcean | Pricing |