← all AI news
VentureBeat · 02 Oct 2026 · 6 MIN READ

Amazon's free Strands Decider 2B turns agent decisions into a 100ms local call

open-sourceagentssafetydev-tools
AI news briefing cover for October 2, 2026: Amazon's free Strands Decider 2B turns agent decisions into a 100ms local call

106 milliseconds. That's the median time Amazon's new decision model needs to pick an option on a three-year-old gaming GPU. It isn't writing an essay or planning a trip. It picks one item from a list you hand it and tells you how sure it is. A week ago I'd have called that a niche trick. Today it looks like the most useful agent primitive anyone has shipped this autumn.

The rest of the day was about friction of a different kind: guardrails that trip over ordinary engineering work, a lab tightening its grip on what its safety people say outside the building, and a judge telling publishers that hoping for Google traffic was never a contract.

Strands Decider 2B makes the boring half of an agent free

AWS's Strands Labs released Strands Decider 2B, an open-source model in the genre TypeSafe started with Jev: you give it a fixed set of options, and it returns a probability for each instead of generating text. VentureBeat's write-up has the numbers that matter:

For comparison, Jev itself costs $0.042 per million input tokens through TypeSafe's hosted API. That's already close to nothing, so Amazon's pitch isn't really price. It's that you can run the thing inside your own perimeter, reproduce it, and stop sending every routing decision to someone else's endpoint.

I think Amazon engineer Marc Brooker described the appeal best in TechCrunch's coverage, calling these models

"a workflow step that can be structured in a way that is more reliable, thanks to the confidence scores, thanks to the closed domain of answers."

That's the right framing. Most of what my agents do badly isn't the hard reasoning. It's the hundreds of small forks along the way: which tool, which file, whether to retry, whether to escalate to a human. Running a frontier model on each of those is slow, costs real money, and returns free-form text I then have to parse and hope about. A closed set of answers plus a calibrated score is something I can put a threshold on, and code can branch on a threshold.

This lands one day after OpenAI announced its own Jev-style Decisions API, and TechCrunch notes that "dozens" of similar models have come out of research labs since TypeSafe introduced the idea. So, yes, the genre is crowded. TypeSafe CEO Diogo Almeida wasn't impressed by the newcomers, saying the current batch "seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful." He has a point that 72% on JevBench isn't 99%. VentureBeat also notes that AWS hasn't shown self-hosting actually saves money once you count hardware and ops.

My take: the money was never the point for me. A 2B model that runs on a laptop takes a network hop and a vendor dependency out of every branch in an agent loop. I'm going to put it in front of tool selection in one pipeline this week and log every case where its confidence lands under 0.6. That log will tell me more than the leaderboard does.

Guardrails are failing in both directions

While the decision layer gets cheaper, the safety layer is getting more expensive for the people who just want to ship. VentureBeat talked to developers whose ordinary work keeps getting flagged. An MIT aeronautics student says Claude treats simulated satellite work as suspicious. A robotics builder gets blocked on tasks like building a UI for an arm and SSH-ing into another machine. One founder summed up his experience in a single line:

"They hate the word cyber. Whenever you say cybersecurity, they just flat out, 'Nope.'" — Rohan Balkondekar

To their credit, both labs admit it. OpenAI concedes its safeguards can "slow, pause, or stop legitimate work" and points cyber professionals to its Daybreak Access program. Anthropic says it made "the wrong tradeoff" with Fable 5 and claims Fable 5.1 cuts interventions per session by about 60%, although pentesting and exploit generation still get routed to less capable models. That matters because, according to the same piece, 90% of more than 15,000 surveyed developers now use coding agents every week. A false positive at that scale isn't an annoyance. It pushes people toward open models with no guardrails at all.

Then there's the other direction. OpenAI has cut ties with three safety researchers, according to a WSJ report, saying they "mishandled sensitive information outside established company procedures." They allegedly shared confidential material with an outside AI safety organization. Neither the organization nor the information has been named, and it isn't clear whether the researchers tried internal channels first. The timing is hard to ignore: it comes days after reporting that OpenAI executives brushed off internal safety warnings, and in the same stretch as the agent containment failures and the shelved GPT-6.1 Astra.

Put the two stories side by side and you get an uncomfortable picture. The guardrails users run into are tuned to be jumpy. The ones that would let insiders raise alarms look like they're tuned to be quiet. I'd swap those settings.

Also on the wire

What these share is that the big wins are increasingly coming from small, cheap components: a 2B scorer, a wiki of past mistakes, 16 GPUs. My prediction is that by year end the interesting agent architectures will spend most of their tokens on tiny specialist models and only call a frontier model when the decision layer says it's unsure. The frontier labs should take note, because that's also where their guardrails will matter least.

Source: VentureBeat ↗


Working on something similar?

Say hello — I read every email.