July 2026
UpdateUnverifiedAdded Sep 25, 2026
Pairwise scoring When you compare two experiments in diff mode, you can now record which one produced the better output for each row. Braintrust aggregates these head-to-head judgments into a win rate, so you can run A-vs-B human evaluations without configuring a separate scoring function. See Pairwise scoring for details.
Topics: Governance, Observability
More releases
Every release| Date | Release | Type |
|---|---|---|
| Jul 13 | June 2026unverified Update | Update |
| Jun 12 | May 2026unverified Update | Update |
| Jun 5 | March 2026unverified Update | Update |
| Jun 5 | February 2026unverified Update | Update |
| Sep 11 | September 2026unverified Update | Update |
| Sep 11 | August 2026unverified Update | Update |
| Apr 29 | April 2026unverified Update | Update |
| Jan 10 | January 2026unverified Update | Update |
Also shipped on Jul 22, 2026
in July 2026Sources: each vendor's own release notes, changelogs and GitHub releases, read daily to monthly by how often it posts. Logos via logo.dev; trademarks belong to their owners.