A benchmark released by Microsoft and Hugging Face asks a practical question about AI agents. ThinkingBox tracks agents changing business records through tools. It checks whether requested tasks finish correctly. The benchmark grades the ending system state, not the agent’s wording. A fluent answer or successful tool call alone does not count.
SOC 2 Ready in 14 days. Three sessions from you.

Enterprise buyers will not put your product near their customer data without a SOC 2 report. Sprinto gets you audit ready in 14 days, across three working sessions. AI agents do the collecting, you approve. Your auditor signs off.
The benchmark contains 507 policy conditioned workflows. They span retail, hospitality, auto insurance, neobank information technology, and consulting support. Each task checks final backend state against hidden executable conditions. It also checks for prohibited side effects. Some tasks score final response disclosures, confidentiality, and consistency. The paper evaluates 18 proprietary and open weight models, with 20 trials per task. Each attempt resets the task and starts an isolated tool session. The simulated user can reveal needed details when the agent asks.
This matters because agent evaluations often report whether one attempt succeeded. A single successful run can show that a model found a workable path once. It does not establish repeatability. Intermediate choices and tool sequences can vary between trials. The paper reports pass at one and all 20 runs succeeding.
Claude Opus 5 scored 66.50 percent on pass at one. It succeeded on all twenty runs for 47.53 percent of tasks. Kimi K3 scored 57.37 percent on the first measure. It succeeded on all runs for 17.60 percent of tasks. These are results for those model versions and this task set. They are not a league table for every current AI system.
The gap makes reliability visible in a way that an average score can hide. A system may complete a task on an isolated attempt, then fail across repeated trials. For organizations, that distinction affects how much oversight a workflow needs. The paper’s tests do not establish production error rates. Real interfaces, data, permissions, and task definitions differ.
ThinkingBox also highlights a measurement problem. Many failed trials ended cleanly after state changing actions. The paper reports no explicit final error in those examples. A valid update could still miss a required state change or create an extra effect. A monitor checking only tool activity may misclassify incomplete work as finished. Termination alone can produce the same mistake. A stronger audit compares final records with the requested outcome and checks for unintended changes.
200+ Proven Ways to Make Money With AI in 2026
The next wave of millionaires will be people who figured out how to make AI work for them.
The window to get ahead is still open. But not for long.
Here are 200+ proven ways to make money with AI in 2026.
Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.
That method has limits because the 507 workflows are synthetic and policy conditioned. Researchers built them inside a sandbox, not collected from operating companies. The paper remains a research preprint. The reviewed sources contain no independent reproduction of its results.
The benchmark cannot settle whether an agent is safe for a particular deployment. A real organization must assess access controls, privacy, and security. It must also assess reversibility, escalation, and human review.
The authors’ framing separates a plausible explanation from an accomplished operation. An agent can narrate a completed booking while the record remains unchanged. It can also make an unintended second change after one correct update. Outcome checks make these differences observable. They give evaluators a clearer target than conversational confidence alone.
For readers assessing agent claims, the central question is repeatability, not whether a model completed a task once. Check what was measured and how often each task was repeated. Then ask whether the data matched the requested result. ThinkingBox offers one research framework for that audit. Its scores describe a controlled benchmark, not workplace performance.
Blu Dot surpasses 2,000% ROAS with self-serve CTV ads
Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Here’s how:
After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.
The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.
“For CTV campaigns, Roku has been a top performer,” said Claire Folkestad, Paid Media Strategist, Blu Dot. “Comping to our other platforms, we have seen really strong ROAS… and highly efficient CPMs, lower than any other CTV partner we've worked with.”
Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.




