Skip to content

New serverless models

UpdateUnverifiedAdded Sep 24, 2026

The following models are now available on serverless : zai-org/GLM-5.3-Flash : 1,000,000 context length, FP8 quantization, function calling and structured outputs. Pricing: $0.15 input / $0.50 output / $0.03 cached input (per 1M tokens).

Topics: Pricing

Together AI's release notes

More Together AI platform releases

Every Together AI platform release
DateRelease
Aug 26Batch jobs in the CLIunverified
Update
Aug 27Billing usage API (beta)unverified
Beta
Aug 27File upload progress in the CLI and Python SDKunverified
Update
Aug 27Slurm cluster kubeconfig for project membersunverified
Update
Aug 27Legacy API key regenerationunverified
Update
Aug 27Model deprecationsunverified
Deprecation
Aug 25Model deprecationsunverified
Deprecation
Aug 25New models available for fine-tuningunverified
Update

Also shipped on Aug 26, 2026

Together AI in August 2026

Sources: each vendor's own release notes, changelogs and GitHub releases, read daily to monthly by how often it posts. Logos via logo.dev; trademarks belong to their owners.

New releases by email

Tuesday mornings: the week's data, AI and developer-tools releases, only in weeks when something shipped.

Double opt-in. Unsubscribe any time.