In partnership with

Anthropic published a new Transparency Hub model report on October 2. It summarizes Claude Sonnet 5.5 safety and capability results. It also describes the company’s evaluation program. Anthropic released Sonnet 5.5 on September 28. The pages show what Anthropic tested. They also identify some material its report leaves out.

Stop typing AI prompts. Start talking.

You think 4x faster than you type. So why are you typing prompts? Wispr Flow turns your voice into ready-to-paste text inside any AI tool. Speak naturally, tangents and all, and Flow cleans it up. Available on Mac, Windows, iPhone, and Android.

The report is not an independent certification. Anthropic designed the evaluations, describes the results, and makes the deployment decisions. Its system card says Sonnet 5.5 performed better than its predecessor. The comparison covers many pre deployment evaluations. The card also says it remains below Opus 5.5 in several areas. These are company findings, attributed to the model developer.

One unusually useful disclosure is methodological. Anthropic says the system card omits some evaluations. These require substantial human work for trustworthy results, and were not considered critical to its conclusions. It also says the Sonnet card is shorter than some previous cards. That choice makes the document easier to scan, but limits direct comparison with earlier reports.

The report covers several different questions that should not be collapsed into one label. The card discusses coding and other capability benchmarks. It also covers cyber evaluations, safety protections and risks from model behavior. A score on a defined task measures performance under that task’s conditions. It does not settle reliability across fields. Nor does it establish truthfulness in ordinary use or safety in every deployment.

External benchmark pages add another kind of measurement. They do not reproduce every company test. Vals lists Sonnet 5.5 at 64.14 percent on Terminal Bench 4.0. It reports an interval of plus or minus 1.01 points. The model placed second among 43 entries. That benchmark concerns terminal tasks, not the biological or alignment evaluations described in Anthropic’s card.

The Vals page also notes that some benchmarks use different providers and parameters. Its displayed default provider is Anthropic, with maximum compute effort. A result therefore belongs to a particular test setup. Without matching model versions, prompts, tools and scoring rules, comparisons can mislead. Runtime settings matter too, so separate evaluations are not a controlled head to head test.

200+ Proven Ways to Make Money With AI in 2026

The next wave of millionaires will be people who figured out how to make AI work for them.

The window to get ahead is still open. But not for long.

Here are 200+ proven ways to make money with AI in 2026.

Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.

That distinction is practical for organizations comparing model documentation. A system card can show which risks a developer considered and which tests it chose. It can also describe where safeguards apply. A third party coding benchmark can measure a narrow set of tasks under stated settings. Neither source by itself proves performance in a different workplace or a future model update.

Sonnet 5.5’s card describes safety levels and company assessments. It does not guarantee harmful outputs cannot occur. Pre deployment evaluations are snapshots. User prompts, connected tools, access controls and the model version can change the conditions after release. The available pages do not provide a universal probability of error for every user or task.

A careful reader can ask what a metric counts and who ran the test. They can check the model version and what was excluded. Anthropic’s new hub makes some answers easier to locate. It presents the company’s own explanations. Its stated omissions show why transparency is useful without being complete.

Before comparing scores, check the exact model build and system prompt. Also compare tool access, sampling settings, retries and scoring rules. A leaderboard position reflects only entries captured in that benchmark’s update. It cannot serve as a general ranking beyond the measured task. Today’s AI check in is about interpreting evidence, not declaring a model safe or unsafe. We will continue separating a developer’s assessment from independent measurements and deployment outcomes.

How Jennifer Aniston’s LolaVie brand grew sales 40% with CTV ads

For its first CTV campaign, Jennifer Aniston’s DTC haircare brand LolaVie had a few non-negotiables. The campaign had to be simple. It had to demonstrate measurable impact. And it had to be full-funnel.

LolaVie used Roku Ads Manager to test and optimize creatives — reaching millions of potential customers at all stages of their purchase journeys. Roku Ads Manager helped the brand convey LolaVie’s playful voice while helping drive omnichannel sales across both ecommerce and retail touchpoints.

The campaign included an Action Ad overlay that let viewers shop directly from their TVs by clicking OK on their Roku remote. This guided them to the website to buy LolaVie products.

Discover how Roku Ads Manager helped LolaVie drive big sales and customer growth with self-serve TV ads.

The DTC beauty category is crowded. To break through, Jennifer Aniston’s brand LolaVie, worked with Roku Ads Manager to easily set up, test, and optimize CTV ad creatives. The campaign helped drive a big lift in sales and customer growth, helping LolaVie break through in the crowded beauty category.