Jamf deployed a system for automatically limiting spending on Claude models in Amazon Bedrock
Jamf published a production serverless system that tracks individual engineers' daily spending on Amazon Bedrock and gradually blocks more expensive models (Claude Opus at 80 %, Claude Sonnet at 100 % of the budget) without interrupting ongoing work.
Jamf, which manages security for Apple devices for more than 76 000 organizations, described a production system for controlling spending on Amazon Bedrock on the AWS Machine Learning blog. After giving its engineers broad access to AI-assisted development, the company needed to address the so-called tokenomics problem – AI spending grows according to user behavior rather than preallocated capacity, and a single engineer in an agentic loop can use more tokens in a few hours than a team does in a week.
The system tracks each engineer's daily spending on Amazon Bedrock and gradually restricts access to more expensive models as they approach their budget: at 80 % of the daily budget, it blocks access to Anthropic Claude Opus; at 100 %, it blocks access to Anthropic Claude Sonnet. Access to Anthropic Claude Haiku remains available so engineers can continue working. According to the company, the restrictions take effect within minutes, do not require a new login, and do not disrupt active sessions; they are automatically lifted after the daily reset. For engineers with a legitimate need for higher limits, an exception process through a Slack command grants a temporary limit increase with a record of who approved it and why.
The architecture consists of three serverless components: spend measurement using a view in Amazon Athena over Bedrock logs (converting tokens to dollars based on published model rates), decision-making and notifications, and enforcement through IAM Customer Managed Policies, which AWS Lambda updates every 15 minutes. The design is idempotent – each run recalculates the full list of restricted users from cumulative daily spending, so a missed or repeated run does not cause inconsistencies. As a precaution, an unrecognized model is priced at the highest rate so it cannot bypass the restrictions.
According to the company, the solution is already running in production, and the code is available as an open-source pattern on GitHub (aws-samples/sample-bedrock-spend-enforcement). The rest of the article, including further operational insights, was not part of the available text.
Why it matters
This is a concrete, publicly shared pattern for companies giving employees broad access to expensive AI models – it addresses the tension between cost control and productivity by gradually redirecting users to a cheaper model instead of blocking access outright, while allowing quick, auditable exceptions. It is relevant to FinOps and platform teams managing Amazon Bedrock or similar multi-model platforms with per-user billing.
Relevant practical impact
What this means
For a business
Companies operating generative AI at scale gain a documented production pattern for controlling spending per user without manual approval or interruptions to work – addressing the specific risk of excessive costs when using models in agentic workflows.
Risks and complianceCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.