Grok Bot Is Everywhere. What Does a Finished Task Actually Cost?
xAI's always-on agents are the closest thing yet to what people meant when they said agent. They are also the first product where ordinary users can watch the meter, and the meter is telling a better story than the launch videos. Four bots held a seven-hour meeting with no humans in it and billed the man who asked them to stop. On why a $230 billion company is buying the name of a free product as a search keyword, the wrong number, the six numbers that matter, why routing beats model choice, the tool I use when I want agents on my own models, and the test I would put to every agent that follows this one.

The question is no longer whether an agent can do the job. It is what one finished job costs, and almost nobody selling one will tell you.
The launch
The picture at the top is what the launch looks like. A ring light, a phone on a tripod, a row of robots each with a green tick, every number on the wall pointing upwards, and a mug that says AI creates freedom. I made it with AI, naturally, and it took the brief a little too literally, because look at the panel in the top right. Bot credits at twelve per cent, down sixty-eight, credits running low, consider upgrading. Even the hype cannot draw itself without the meter creeping into the corner of the frame.
Grok Bot is suddenly everywhere. My feed has become a parade of people demonstrating assistants that clear their inbox, work their pipeline, book their travel and keep going after they have closed the laptop and gone to bed. There is clearly a sizeable marketing budget behind it, and the affiliate links are doing what affiliate links do.
Strip that away and there is something interesting underneath. xAI is not describing a chatbot. Its own words are “your team of always-on agents” that “have their own computer, work inside tools and apps like you do, and keep working 24/7”. They sign into the software you already use, move between applications, and bring a job back only when it needs your approval. It launched in beta on 11 August and comes bundled into the SuperGrok tiers and, through the Cursor tie-up, into the Cursor plans as well. If you want to try it yourself, the front door is x.ai/bot.
For two years we have been told agents are just around the corner, and mostly what turned up was a chat window with a plug-in. This is closer to the thing people actually meant. You give it a job rather than a prompt. It works across software rather than inside a box. It carries on without you.
So I am not going to ask whether Grok Bot is clever. Plenty of people are on that already, and the answer is clearly “often”. The more useful question is the one that never appears in a launch video, which is whether it makes economic sense.
Buying the front door
There is a second thing the marketing push tells you, and you can see it by typing the name of the competitor into Google.

OpenClaw is the thing Grok Bot is chasing. It started in November 2025 as one Austrian developer’s side project, was renamed twice in a single week in January after a trademark complaint from Anthropic, and passed a quarter of a million GitHub stars by March, which makes it one of the fastest-growing open-source projects there has ever been. It runs on your own machine, talks to you over WhatsApp, Telegram or Signal, and does the same category of work: the inbox, the calendar, the browser, the errands. You bring your own API key or point it at a local model. There is no subscription, no hosted tier, and no enterprise edition, and it is stewarded by a non-profit foundation whose donors include OpenAI, Amazon and Nvidia, none of whom control it. Its tagline is “nobody’s business model”. Version 2.0 shipped on 30 August, nineteen days after Grok Bot launched.
So the incumbent in the category of agents that actually do things on a computer is free, and the challenger is a company that raised twenty billion dollars in January at a valuation of two hundred and thirty billion, and now sits inside SpaceX at a combined one and a quarter trillion. It cannot beat free on price. What it can do is buy the front door. Search for the free product and the first screen you see is the paid one, with six sitelinks and a boast about a million visits a month, and the free product’s own website is the first thing below the fold. I do not know what xAI is spending on search advertising. I do know what the top of the page looks like, and “nobody’s business model” is a fine tagline right up until somebody with a business model buys your name as a keyword.
This matters for the essay because the two products have opposite meters. With OpenClaw you pay the model provider directly, every token is itemised on an API bill you can read, and you decide which model handles which step, down to running the routine ones on your own hardware for nothing. With Grok Bot the meter is inside the product, the allowance is unpublished, and the model is chosen for you. One of those lets you calculate a cost per task on a Tuesday afternoon. The other one, as its own bot told me below, does not. That open, own-models, own-meter side of the market is the one I think is useful, and it is the side ATLAS sits on: the governor that decides which model gets which step, holds the budget before a token is spent, and keeps the ledger, so the meter is not just visible but yours. The most expensive word on the internet this month may well be the name of a free product.
The wrong number
The number most people reach for first is the subscription, and that is the wrong number. I say this as someone who once wrote a whole essay after a quota meter reset under me while I was doing the washing up. Price is not cost. The ticket is not the fuel.
The cost of an autonomous agent is not thirty dollars a month, or two hundred, or whatever the tier happens to be that unlocks it. The cost is the cost per successfully finished task, and the two numbers can be very far apart, because an agent does not consume intelligence in neat units. A single instruction can set off planning, a browser session, a dozen model calls, tool calls, a failed action, a retry, a self-check and a conversation between agents before you see anything at all. The interface shows you one task. The meter sees the whole expedition.
xAI has half admitted this. Grok Bot “comes with its own usage, separate from your Grok and Cursor plans”. Read that as a product decision and it is sensible. Read it as a disclosure and it is more interesting. Autonomous work has its own cost profile, and the company that built the product knows it.
So I asked the bot. Spelling mistake included, which it handled without comment, which is more than most humans manage.

The answer is better than most of the marketing. There is no separate fee. You are billed “by how much work a task does (steps and tokens), not per message, so a quick draft is cheap and a long research job burns more”. And then, unprompted, the line I would have paid for: the docs “don’t publish a flat ’$ per task’, so that’s the honest meter”. The product’s own agent has just told you that the number you need is the one nobody publishes, and pointed you at a usage screen to work it out yourself. I have had less candid conversations with vendors’ sales directors.
What the meter says
The useful thing about a product this widely used is that people post their meters. None of it is a benchmark. All of it is more informative than the pricing page, which does not publish the size of the weekly allowance on any tier. This is what the first six weeks have produced.
| Reported | What happened |
|---|---|
| Day one | A six-agent “business” consumed 42% of its weekly allowance on the day it was set up. |
| Week two | About a hundred basic chat completions plus one ten-minute script used 5% of a week. The author called this “truly terrible” for real business use. |
| Week two | A SuperGrok Heavy subscriber reverse-engineered their allowance at roughly 16.4 million tokens and noted that “system-prompt overhead eats into every turn”. |
| September | A user who had run a week without trouble hit the whole weekly limit in a single day. His working theory was that the bot loads the entire history of every chat on every turn. |
| September | A user asked his QC, UAT, coding and design bots to stay quiet and report only to the command agent. They kept reviewing each other’s work anyway. The weekly meter hit 100% and the account locked. |
| September | A related report of agents “talking to each other for seven hours straight” with no human in the loop, on on-demand billing. |
The moderator’s explanation of the locked account deserves framing. “Each bot-to-bot message runs a turn that counts toward your weekly usage.” Somewhere in a data centre, four agents held a seven-hour meeting with no agenda, no humans and no biscuits, and billed it to the man who had asked them to stop. That is not artificial intelligence. That is middle management.
Two structural facts sit underneath all of this. The allowance is weekly and resets on a Monday, so one heavy Tuesday can push you onto the meter for the rest of the week. Once it does, on-demand is billed from token cost at the underlying model’s price, with only the account-level monthly cap standing between you and the bill, and, until a recent feature request is acted on, no warning that you have crossed over. One early tester summed up his month by saying he had used fewer tokens in the previous five years than in the last thirty days. That is not a complaint about capability. It is a complaint about a meter nobody can see.
I have some sympathy. I have a weekly allowance now. The last time I had one of those I was at school, and I spent it on the same thing, which was something that talked back.
The invisible work
Part of this is an interface problem, and it is one I recognise from managing people. When I ask someone to check these leads and tell me which five are worth calling, I know there is work hidden inside the sentence. They have to read, search, compare, reject, verify a couple of things, form a judgement and then explain it. I do not see any of that. I see five names and I judge them on it.
Agents hide the intermediate steps at least as well as a good analyst does, and that is one of their strengths. It also obscures their economics completely. A polished demo shows you prompt, then finished task. The system underneath looks more like prompt, plan, search, model call, browser action, model call, failed click, retry, another model call, validation, and then, at last, an answer. If all I see is the two ends, I have no idea what the middle cost, and the middle is where the money went.
The user who noticed his bot loading its entire history on every turn had put his finger on something I wrote about last week without knowing a product would prove the point this quickly. What the system knows when the model starts thinking is an engineering decision. Reload everything, every time, and you are paying to re-describe the world on every turn. That is not intelligence. It is a very expensive habit.
Which is why token counts are far more interesting in agentic systems than they ever were in chat, and also why tokens alone are no longer enough. What a business needs is closer to a manufacturing unit cost. I spent long enough around companies that made physical things to know that a factory reporting its output as units attempted would not survive the quarter.
The six numbers
If someone shows me an agent replacing administrative work, I want six numbers beside the demo. None of them is complicated. They are simply never provided.
The first is tasks attempted, meaning how many pieces of real work we gave it. Jobs, not prompts. The second is tasks completed without intervention, and completed means finished, correctly, without a person taking over halfway. It does not mean the model produced something. The third is human correction time, because if the agent saves thirty minutes and needs fifteen minutes of checking it saved fifteen, and if the checking is done by the most expensive person in the building it may have saved nothing at all.
The fourth is total agent cost: the allowance it consumed, the on-demand spend, the supporting services, and the share of the subscription that was really buying this. The fifth is cost per accepted task, which is everything above divided by the number of tasks you actually kept. This is the number that finally lets you compare an agent with a person, with conventional automation, or with not doing the work at all.
The sixth is failure consequence, and it belongs on the page for anything that matters. A wrong meeting summary costs nothing. A wrong invoice payment, customer email, production change or compliance decision costs something, and that something sits on the same page as the savings or the page is fiction. The model is not the system, and the cost of an agent cannot be separated from the cost of it being wrong.
Worked through with numbers I have made up for the arithmetic rather than measured, it looks like this.
| Tasks given to the agent in a week | 100 |
| Finished cleanly, no hands on | 70 |
| Finished after a human fixed something | 20 |
| Abandoned and done by a person | 10 |
| Human correction time | 20 tasks × 10 min ≈ 3.3 hours |
| Agent cost for the week | allowance share + on-demand, say $140 |
| Human time at a loaded $80 an hour | ≈ $265 |
| Cost per accepted task | ($140 + $265) ÷ 90 ≈ $4.50 |
Now do the sum for the process it replaced. If a person was doing those ninety tasks at four minutes each, that is six hours, or about $480, or $5.30 a task. The agent wins, narrowly, and only if the ten abandoned tasks cost nothing on their way to being abandoned. Change the correction time to twenty minutes and the person wins. That is the whole point. The answer lives in the denominator, and the denominator is accepted work rather than attempted work.
The builders’ version of this, five numbers measured per decision inside the system, is in the context essay. This is the buyers’ version. The two should agree with each other, and when they do not, someone is being optimistic.
The illustration at the top, incidentally, has its own dashboard: 1,248 tasks completed, 37 hours of runtime, estimated cost $87.42. That is seven cents a task. If it were true nobody would be writing to the Cursor forum. The picture was generated by a model that has read a great many launch pages, and it drew the number those pages implied.
Routing may matter more than the model
One piece of user feedback interested me more than the rest. Several people who looked at what their bots were actually doing complained that simple work was being sent through the heavyweight model tier. When the routing changed and cheaper models took over, the same people reported the bots became, in one user’s phrase, “quite stupid at times”. There is no model picker. Billing follows whichever model served the turn.
I cannot verify anyone’s measurements, but the architectural point is real, and it is the thing I would fix first if I owned the product. An autonomous system should not use frontier intelligence for everything. Checking whether an email contains an invoice does not require the same brain as deciding whether a contract creates an unusual commercial risk. The efficient shape is boring and I have written it before. Code where certainty is enough, small models where specialisation is enough, frontier models where genuine reasoning is required, and escalation only when the cheaper tier admits defeat.
That is the difference between an AI agent and an expensive sequence of model calls wearing an assistant’s face. The people already routing their repeatable steps through a plain automation tool and keeping the agent for judgement have understood this faster than the product has.
It is also, for what it is worth, the problem I built ATLAS to solve for my own work. ATLAS is the productised version of The Governor, the agentic operating system I wrote about in July, and it is what I use internally when I want agents running against my own models rather than somebody else’s meter. Routing starts local and escalates to a frontier model only when complexity, risk and earned trust justify it. The budget is reserved as a hold before a single token is generated, rather than checked afterwards, because a check loses to two requests arriving at once. Something that did not do the work decides whether the work was done, and a refuted run is rolled back rather than kept. Every routing choice lands in a ledger joined to the money, so the question at the end of the week is not what did this cost but who spent it, on what, and what were they told. If you want to run your own models, that is the tool. If you want Grok Bot, the same questions apply. You simply cannot answer them yet.
The product may still be very good
The meter table above reads like a prosecution and it is not meant as one. Nothing here means Grok Bot is a bad product. The description xAI gives, persistent agents with their own computer, memory and access to ordinary applications, is far closer to the architecture businesses have been waiting for than another conversational assistant with a plug-in menu. The people complaining loudest about the allowance are, almost without exception, people who are still using it. They describe it as useful for personal administration and repetitive knowledge work in the same post as the complaint about running out.
That combination is worth noticing. The product may already be useful enough that the biggest grievance is not capability. It is wanting more of it at a price you can plan around. That is a good problem for a young product to have. It is still a problem, and it is a problem with a meter attached.
The hype test
So here is the test I would apply to Grok Bot, and to every autonomous agent that follows it, and there will be a lot of them. Do not show me that it can do the task once. Show me a hundred real tasks. Show me the percentage finished without intervention, the human time spent supervising, the total usage consumed, the failures and what they cost, and the resulting cost per accepted outcome. Then put the existing process beside it and let me look at both.
If those numbers work, the hype is justified and I will be first in the queue. If they do not, we have not automated the job. We have automated the demonstration.
Where this gets interesting
The important change is not Grok Bot. It is that autonomous computer agents have become accessible enough that ordinary businesses can start collecting those numbers for themselves, on real work, this quarter, without a consultancy and without a pilot programme that MIT will later count among the 95% that showed no bottom-line impact.
For two years the question has been whether AI can do this. The answer has arrived and it is mostly yes. The next question is more useful and much harder. Can it do this reliably, repeatedly and cheaply enough that I should redesign the business around it? Grok Bot may turn out to be one of the products that helps us answer that, precisely because it is the first one where the meter is visible enough to argue about.
Just do not confuse a spectacular demo with the answer. And if your bots start holding meetings without you, check the bill before you check the minutes.
Thanks for reading

One last thing, for fun. Both pictures in this essay were generated by AI, and in an essay about what AI work actually costs it seemed only fair to show you the seconds. Look closely at the close-up. The mug’s reflection on the desk has spelt freedom with a V and a K. Someone has handwritten a thank-you note on a sheet of glass that is hovering in mid-air with nothing holding it up. The book spines read like the shopping list of a motivational poster. Back in the wide shot, the letters AI are floating over the skyline outside the window, and every arrow in the building points up except the one on the bill. I make it five, and I expect you will find more.
It is still considerably better than I can draw, which is rather the point of the whole essay. The demo is spectacular. It is still worth checking the work.