Compare
FlockTab vs provider spend limits and AI gateways
OpenAI, Anthropic and Google all let you cap what you spend. Gateways such as LiteLLM, Portkey, Helicone, Cloudflare AI Gateway and OpenRouter add budgets across providers. Many teams use one already. Here is what each one does, from its own docs, and where FlockTab is different.
We read every source on 6 October 2026, and each claim links to it. If something has changed, tell us and we will fix it.
The short answer
- Turn on your provider's hard limit. It costs nothing, and it is your last line of defence.
- If you serve an app to many users, a gateway gives you budgets per key or per user, plus caching, routing and analytics.
- If you run coding agents such as Claude Code, Codex or Gemini CLI, FlockTab gives each one its own tab, with a hard cap and a kill switch. On Team and Fleet, Agent Guard holds a risky tool call until you allow it. Agents that log in with a Claude or ChatGPT plan run under the same gate and kill switch.
Side by side
From each product's own docs. “Not described” means we found nothing in the docs, not that it cannot be done.
| FlockTab | OpenAI | Anthropic | Google Gemini API | LiteLLM | Portkey | Helicone | Cloudflare AI Gateway | OpenRouter | |
|---|---|---|---|---|---|---|---|---|---|
| A money cap that blocks calls | Yes, per agent[34] | Per organisation and project[1] | Per organisation and workspace[4] | Per project (experimental) and billing account[8] | Yes[14] | Enterprise and some Pro plans[18] | Only as a cost rate limit[22] | Yes[26] | Per key; workspace budgets on Enterprise[31][30] |
| Holds the estimated cost before each call | Yes[34] | No: the limit is not instant[1] | Not described[4] | No: about 10 minutes of lag[8] | Yes, by default[14] | Not described[18] | Not described[22] | No: checks recorded spend[26] | No: checks recorded spend[30] |
| A budget per key or per agent | Per agent[34] | Not described[1] | Per workspace; per user only for Claude Code[5] | No: keys share the project's limits[8] | Key, user, team, customer, agent[14] | Per API key[17] | Rate limit per key or user[22] | By user or agent metadata, not per token[26][27] | Per key and per member[32] |
| Money per short window (minutes or hours) | Hourly caps[35] | Monthly[1] | Monthly[4] | Per 10 minutes, set by Google per tier[9] | Any window[14] | A day at the shortest[19] | Per window of 60 seconds or more[22] | Rolling or fixed; sizes not listed[26] | A day at the shortest[30] |
| A person approves a tool call before it runs | Agent Guard, on your machine (Team and Fleet)[37] | Remote MCP tools only[3] | Managed Agents (beta) server tools[7] | Computer use safety check[10] | Rules only, no person[15] | Rules on MCP tool arguments[20] | Not described[24] | Not described[26] | Not described[32] |
| One cap across providers | Yes, per agent[34] | OpenAI only[1] | Anthropic only[4] | Google only[8] | Yes[14] | Not stated[19] | Rate limit across the gateway[22] | Yes[26] | Yes; your own keys only if counted[32] |
Provider limits
OpenAI
Anthropic
AI gateways
LiteLLM
Good at
The most complete budgets here: per key, user, team, customer and agent, with any reset window and several windows on one key.[14] Before each call it holds the estimated cost, and it rejects the call if the budget can't cover it.[14] It can block tool calls by rule, and you can run it yourself for free.[15][16]
Portkey (now Prisma AIRS AI Gateway)
Helicone
Cloudflare AI Gateway
What FlockTab does differently
A tab for each agent
Each coding agent gets its own tab, with a hard cap and a window: one run, an hour, a day, a week, a month, or no end.[35] Before a call reaches the provider, FlockTab estimates its cost from the model and max_tokens and holds it against the tab. After the reply, the call settles at the tokens the provider reported. If the ledger can't be reached, the call is refused.[34] LiteLLM holds an estimate too. Cloudflare and OpenRouter check spend already recorded, so a burst can go over.[14][26][30]
Agent Guard
Before a tool call runs in Claude Code, Codex or 8 other harnesses, FlockTab checks it.[38] A force-push, a production deploy, a payment or a recursive delete is stopped by rule. Anything else that can change something is judged, and the reason is written in plain words.[36] A flagged call waits for a person: allow it in time and it runs, otherwise it is refused.[37] These are the tool calls your agent runs on your own machine. The gateways here describe tool rules, but no person approving a call. OpenAI and Anthropic ask for approval only for tools that run on their own servers.[3][7] Agent Guard is on Team and Fleet, and off for each agent until you turn it on.[37]
Plan logins, too
Claude Pro or Max and ChatGPT plan logins run under the same gate and kill switch. Their calls are recorded at list price and never charged to the tab.[39][38]
One switch
Close one tab, and that agent's next call is refused with the reason.[34] Close all tabs, and every agent stops.[33]
When FlockTab is not the right fit
If you serve an app to many end users and need caching, fallbacks or prompt analytics, a gateway is built for that. If you call one provider from one script, that provider's hard limit may be all you need.
Solo is free for 3 named Agents. See pricing
Sources
Each page was read on 6 October 2026.
- [1]OpenAI: Spend limits
- [2]OpenAI: Rate limits
- [3]OpenAI: MCP servers
- [4]Anthropic: Rate limits
- [5]Anthropic: Workspaces
- [6]Anthropic: Creating and managing Workspaces in the Claude Console
- [7]Anthropic: Permission policies (Managed Agents)
- [8]Google: Billing, Gemini API
- [9]Google: Rate limits, Gemini API
- [10]Google: Computer use, Gemini API
- [11]Google Cloud: Create, edit, or delete budgets and budget alerts
- [12]Google Cloud: Manage spend cap budgets
- [13]Google Cloud: Cloud Billing release notes
- [14]LiteLLM: Budgets, Rate Limits
- [15]LiteLLM: Tool Permission Guardrail
- [16]LiteLLM: Enterprise
- [17]Portkey: Enforce Budget Limits and Rate Limits for Your API Keys
- [18]Portkey: Budget Limits
- [19]Portkey: Usage & Rate Limit Policies
- [20]Portkey: Guardrails (MCP Gateway)
- [21]Portkey: Guardrails
- [22]Helicone: Custom LLM Rate Limits
- [23]Helicone: Alerts
- [24]Helicone: Sessions
- [25]Helicone: Self-Hosting Helicone
- [26]Cloudflare: Spend limits, AI Gateway
- [27]Cloudflare: Authenticated Gateway, AI Gateway
- [28]Cloudflare: Caching, AI Gateway
- [29]Cloudflare: Pricing, AI Gateway
- [30]OpenRouter: Workspace Budgets
- [31]OpenRouter: Limits
- [32]OpenRouter: Spend Controls
- [33]FlockTab: Home: the kill switch
- [34]FlockTab: How it works
- [35]FlockTab: tab CLI reference
- [36]FlockTab: Agent Guard
- [37]FlockTab: Agent Guard docs
- [38]FlockTab: Supported harnesses
- [39]FlockTab: Your own logins
- [40]FlockTab: Outside spend
- [41]FlockTab: Outside spend docs