Compare

FlockTab vs provider spend limits and AI gateways

OpenAI, Anthropic and Google all let you cap what you spend. Gateways such as LiteLLM, Portkey, Helicone, Cloudflare AI Gateway and OpenRouter add budgets across providers. Many teams use one already. Here is what each one does, from its own docs, and where FlockTab is different.

We read every source on 6 October 2026, and each claim links to it. If something has changed, tell us and we will fix it.

The short answer

  • Turn on your provider's hard limit. It costs nothing, and it is your last line of defence.
  • If you serve an app to many users, a gateway gives you budgets per key or per user, plus caching, routing and analytics.
  • If you run coding agents such as Claude Code, Codex or Gemini CLI, FlockTab gives each one its own tab, with a hard cap and a kill switch. On Team and Fleet, Agent Guard holds a risky tool call until you allow it. Agents that log in with a Claude or ChatGPT plan run under the same gate and kill switch.

Side by side

From each product's own docs. “Not described” means we found nothing in the docs, not that it cannot be done.

FlockTabOpenAIAnthropicGoogle Gemini APILiteLLMPortkeyHeliconeCloudflare AI GatewayOpenRouter
A money cap that blocks callsYes, per agent[34]Per organisation and project[1]Per organisation and workspace[4]Per project (experimental) and billing account[8]Yes[14]Enterprise and some Pro plans[18]Only as a cost rate limit[22]Yes[26]Per key; workspace budgets on Enterprise[31][30]
Holds the estimated cost before each callYes[34]No: the limit is not instant[1]Not described[4]No: about 10 minutes of lag[8]Yes, by default[14]Not described[18]Not described[22]No: checks recorded spend[26]No: checks recorded spend[30]
A budget per key or per agentPer agent[34]Not described[1]Per workspace; per user only for Claude Code[5]No: keys share the project's limits[8]Key, user, team, customer, agent[14]Per API key[17]Rate limit per key or user[22]By user or agent metadata, not per token[26][27]Per key and per member[32]
Money per short window (minutes or hours)Hourly caps[35]Monthly[1]Monthly[4]Per 10 minutes, set by Google per tier[9]Any window[14]A day at the shortest[19]Per window of 60 seconds or more[22]Rolling or fixed; sizes not listed[26]A day at the shortest[30]
A person approves a tool call before it runsAgent Guard, on your machine (Team and Fleet)[37]Remote MCP tools only[3]Managed Agents (beta) server tools[7]Computer use safety check[10]Rules only, no person[15]Rules on MCP tool arguments[20]Not described[24]Not described[26]Not described[32]
One cap across providersYes, per agent[34]OpenAI only[1]Anthropic only[4]Google only[8]Yes[14]Not stated[19]Rate limit across the gateway[22]Yes[26]Yes; your own keys only if counted[32]

Provider limits

OpenAI

Good at

Hard monthly limits for the whole organisation and for each project. A request over a limit fails with an error that names the limit.[1] Rate limits can be set for each project.[2]

Keep in mind

The limit is not instant, so spend can go a little over.[1] The docs describe no limit per API key, and the limits cover OpenAI traffic only.[1]

Anthropic

Good at

A hard monthly cap for each usage tier, and your own lower limit on top.[4] Each workspace can have its own spend limit and email alerts. The Claude Code workspace can cap each user.[5][6]

Keep in mind

The docs don't say how fast a limit takes effect. The default workspace cannot be capped.[4] Keys belong to a workspace and have no limit of their own. It covers Anthropic only.[6]

Google

Good at

The Gemini API has a monthly cap for each project and a cap for each billing account.[8] It also limits spend per 10 minutes, a brake the other providers here don't document.[9]

Keep in mind

The project cap is experimental and can run about 10 minutes over.[8] Google Cloud budgets send alerts but never stop spend.[11] The newer spend caps for Vertex AI are in preview: one project and one service each, and not instant.[12][13]

AI gateways

LiteLLM

Good at

The most complete budgets here: per key, user, team, customer and agent, with any reset window and several windows on one key.[14] Before each call it holds the estimated cost, and it rejects the call if the budget can't cover it.[14] It can block tool calls by rule, and you can run it yourself for free.[15][16]

Keep in mind

We found no step where a person approves a single tool call.[15] Budgets need a database. Without one, the proxy-wide budget lets calls through.[14] Calls with no token price, and batch jobs, are checked against recorded spend instead.[14]

Portkey (now Prisma AIRS AI Gateway)

Good at

Guardrails on prompts and responses, and checks on MCP tool arguments before the tool runs.[21][20] Budgets on API keys, and usage limits by key, user, model or provider.[17][19]

Keep in mind

Its docs now name it Prisma AIRS AI Gateway.[21] Budgets are on Enterprise and some Pro plans, and the limits engine needs self-hosted Enterprise.[18][19] The docs don't say whether a budget is checked before or after a call.[18]

Helicone

Good at

Seeing what your agents do: sessions, users and cost per request.[24] A cost limit per time window for each user, and you can run it yourself.[22][25]

Keep in mind

There is no budget feature as such. The cost limit is set by a header the client sends.[22] Cost alerts tell you, but don't block.[23]

Cloudflare AI Gateway

Good at

Caching of identical requests, and core features free on every plan.[28][29] Spend limits per user, agent, model or provider. Over a limit, it can switch to a cheaper model instead of blocking.[26]

Keep in mind

Limits are checked against spend already recorded, so a burst of calls can go over.[26] API tokens are account-wide, so there is no budget per token.[27]

OpenRouter

Good at

Each API key can have a spending cap.[31] A member's budget adds up across all their keys, and a workspace can pool one cap across every member and key.[32]

Keep in mind

Budgets are checked before routing, but against recorded spend, so calls already running can go over.[30] Workspace budgets are Enterprise, and the shortest reset is a day.[30] Spend on your own provider keys doesn't count unless you turn that on.[32]

What FlockTab does differently

A tab for each agent

Each coding agent gets its own tab, with a hard cap and a window: one run, an hour, a day, a week, a month, or no end.[35] Before a call reaches the provider, FlockTab estimates its cost from the model and max_tokens and holds it against the tab. After the reply, the call settles at the tokens the provider reported. If the ledger can't be reached, the call is refused.[34] LiteLLM holds an estimate too. Cloudflare and OpenRouter check spend already recorded, so a burst can go over.[14][26][30]

Agent Guard

Before a tool call runs in Claude Code, Codex or 8 other harnesses, FlockTab checks it.[38] A force-push, a production deploy, a payment or a recursive delete is stopped by rule. Anything else that can change something is judged, and the reason is written in plain words.[36] A flagged call waits for a person: allow it in time and it runs, otherwise it is refused.[37] These are the tool calls your agent runs on your own machine. The gateways here describe tool rules, but no person approving a call. OpenAI and Anthropic ask for approval only for tools that run on their own servers.[3][7] Agent Guard is on Team and Fleet, and off for each agent until you turn it on.[37]

Plan logins, too

Claude Pro or Max and ChatGPT plan logins run under the same gate and kill switch. Their calls are recorded at list price and never charged to the tab.[39][38]

One switch

Close one tab, and that agent's next call is refused with the reason.[34] Close all tabs, and every agent stops.[33]

Speed and model limits

Each agent can have a limit on calls per minute (a new agent starts at 120) and a list of allowed models. A call over either is refused before the provider.[35][34] Gateways have rate limits too. Here they belong to each agent.

One ledger

Every decision is one ledger line.[34] On Team and Fleet, connectors read the bills that don't go through a proxy, such as OpenRouter, Cloudflare and GitHub, line by line, and put them beside your tabs.[40][41]

When FlockTab is not the right fit

If you serve an app to many end users and need caching, fallbacks or prompt analytics, a gateway is built for that. If you call one provider from one script, that provider's hard limit may be all you need.

Solo is free for 3 named Agents. See pricing

Sources

Each page was read on 6 October 2026.

  1. [1]OpenAI: Spend limits
  2. [2]OpenAI: Rate limits
  3. [3]OpenAI: MCP servers
  4. [4]Anthropic: Rate limits
  5. [5]Anthropic: Workspaces
  6. [6]Anthropic: Creating and managing Workspaces in the Claude Console
  7. [7]Anthropic: Permission policies (Managed Agents)
  8. [8]Google: Billing, Gemini API
  9. [9]Google: Rate limits, Gemini API
  10. [10]Google: Computer use, Gemini API
  11. [11]Google Cloud: Create, edit, or delete budgets and budget alerts
  12. [12]Google Cloud: Manage spend cap budgets
  13. [13]Google Cloud: Cloud Billing release notes
  14. [14]LiteLLM: Budgets, Rate Limits
  15. [15]LiteLLM: Tool Permission Guardrail
  16. [16]LiteLLM: Enterprise
  17. [17]Portkey: Enforce Budget Limits and Rate Limits for Your API Keys
  18. [18]Portkey: Budget Limits
  19. [19]Portkey: Usage & Rate Limit Policies
  20. [20]Portkey: Guardrails (MCP Gateway)
  21. [21]Portkey: Guardrails
  22. [22]Helicone: Custom LLM Rate Limits
  23. [23]Helicone: Alerts
  24. [24]Helicone: Sessions
  25. [25]Helicone: Self-Hosting Helicone
  26. [26]Cloudflare: Spend limits, AI Gateway
  27. [27]Cloudflare: Authenticated Gateway, AI Gateway
  28. [28]Cloudflare: Caching, AI Gateway
  29. [29]Cloudflare: Pricing, AI Gateway
  30. [30]OpenRouter: Workspace Budgets
  31. [31]OpenRouter: Limits
  32. [32]OpenRouter: Spend Controls
  33. [33]FlockTab: Home: the kill switch
  34. [34]FlockTab: How it works
  35. [35]FlockTab: tab CLI reference
  36. [36]FlockTab: Agent Guard
  37. [37]FlockTab: Agent Guard docs
  38. [38]FlockTab: Supported harnesses
  39. [39]FlockTab: Your own logins
  40. [40]FlockTab: Outside spend
  41. [41]FlockTab: Outside spend docs