What Astra Actually Changed
OpenAI released GPT-6 Astra on September 3, 2026. Within about twenty four hours, a list of nine Astra prompts was circulating everywhere. Negotiate your bills. Watch Facebook Marketplace for underpriced listings. Run nightly QA on your app using a real phone. Spy on your competitors.
We went looking for anybody who had actually run those prompts and published the result. We found nobody. Not one source we could reach showed a completed run of any of the nine. The list went viral on how good it sounded, not on what it did.
That is worth knowing before you spend money. But it would be a mistake to conclude that Astra is nothing. Astra is a real jump, in one very specific direction, and that direction is genuinely useful to a small business. It is just not the direction most of the launch week content claimed.
Astra is not smarter. It has better hands.
The thing Astra does that previous models could not do reliably is operate a computer. It clicks through software you already use. It drives a browser. It works inside a desktop application. On the standard test for this, called OSWorld 2.0, Astra scores 72.6 percent against 65.7 percent for the previous OpenAI model, and it does it in roughly 47 percent less time per task. About forty minutes per task instead of seventy five.
Read that number honestly, though, because two things about it matter. First, that score is a partial score, meaning credit for tasks it got part of the way through. The strict completion rate across the whole field of models is well under half. Second, OSWorld runs in controlled test environments, not on live websites with popups, cookie banners, and login walls. Real websites are harder than the benchmark.
On reasoning, Astra is not the leader. On the hardest reasoning benchmark going, Astra scores 57.2 percent while Claude Fable 5.1 scores 65.0 percent. That comparison appears on OpenAI's own launch page, which is to their credit. They published a row they lost.
The one sentence version: if your bottleneck is thinking, Astra is not your upgrade. If your bottleneck is somebody having to sit and click through the same interface for two hours, it might be.
First, Check If You Can Even Run It
This is the part almost every article skipped, and it is the part that will waste your afternoon if you get it wrong.
OpenAI's launch post said Astra would reach "all ChatGPT Plus, Pro, Business, and Enterprise users." OpenAI's own help center says something narrower. For the Pro plan it lists Astra "in Chat, ChatGPT Work and Codex." For the Plus plan it lists Astra "in ChatGPT Work and Codex." The word Chat is in one sentence and missing from the other.
Plus users noticed. There are two threads on OpenAI's developer forum asking about exactly this, and at the time of writing neither has a reply from OpenAI staff. So we will say what the evidence supports and no more: the help center lists Chat for Pro and omits it for Plus. That is strong but circumstantial. It is not an explicit statement that Plus users are excluded from the chat picker.
How many messages do you get?
OpenAI has not published a number. Their exact wording is that Astra "can use your allowance faster than GPT-5.6 Sol" and that Plus includes "limited Astra usage, with optional credits." That is it.
You will see a figure of five to forty five messages per five hour window quoted confidently in a lot of blog posts. That number traces back to a single post on X, and the person who posted it described it as an estimate. It has since been repeated as fact by sites that never checked. Treat it as a rumor.
| Plan | Price | Astra access | What it means for you |
|---|---|---|---|
| Free | $0 | None | No announced plan to change this. |
| Go | $8/mo | None | Same as Free for this purpose. |
| Plus | $20/mo | Work and Codex | You can try the real capability, in a surface you have to go and find. Small, undisclosed allowance. |
| Pro | $100/mo | Chat, Work, Codex | Five times Plus usage. This is the realistic floor for doing agent work more than once. |
| Pro | $200/mo | Chat, Work, Codex | Twenty times Plus usage. For people running this every day. |
| Business | $25/seat | Limited | Standard seats behave like Plus, with optional credits. |
| API | $10 / $50 per M | Full | No message cap. An uncapped bill instead. Read Section 07 first. |
The honest upgrade advice: if you want to find out whether Astra's computer use does anything for your business, buy one month of Pro at $100 and treat it as an experiment. Plus will not give you enough runs to learn anything. The API will not give you a predictable bill. Run three or four real tasks, decide, and cancel if it did not earn the money back.
The Nine Viral Prompts, Audited
The list originated with a single post on September 4 that showed no results. It has since been republished, uncredited, by at least four different sites presenting it as their own work. One of those sites has "Tested" and "Results" in the headline and contains no output, screenshot, or transcript anywhere in the piece.
Here is what happens when you check each one against what Astra can actually do.
ask for things that do not exist, or carry legal exposure you personally own
are ordinary prompts that work on any model you already pay for
actually use Astra's new capability, and have working demos behind them
The three that do not work
The bill renegotiator
OpenAI's help page lists what its finance feature cannot do, and the list includes "Pay bills," "Move money," and "Change account settings." It states plainly that ChatGPT "cannot take financial actions for you."
The marketplace flipper
Craigslist terms set liquidated damages at $3,000 per day for automated collection. They have collected a $60 million judgment and a $1 million settlement in prior scraping cases. Also, Astra cannot send you a text message.
Nightly QA on a real phone
Astra's computer use runs in a cloud virtual machine. It cannot see or touch a physical device. Driving a real handset needs a device farm, which is an engineering project, not a prompt.
The competitor spy
Signing up under your own name is fine. Signing up under a false one to hide that you are a competitor breaks terms of service almost everywhere, and the risk sits with you, not OpenAI.
The bill renegotiator deserves an extra note, because it is the most repeated of the nine and it has a second problem behind the first. Agent mode "will pause and prompt you to take control of the virtual browser" whenever a task requires a login, and every provider portal requires a login. So even the negotiate half is not automation, it is babysitting.
There is also a disclosure question. California SB 1001 makes it unlawful to use a bot to mislead someone about its artificial identity in a commercial transaction, at $2,500 per violation. And the FTC has already been here: they issued a $193,000 order against DoNotPay, the closest prior product to this idea, partly because the company never tested whether its service performed to the standard it advertised. If you were planning to resell "AI negotiates your bills" as a service, that order is a map of how it goes wrong.
The four that work but do not need Astra
Turn an agency into software. Build a one person company dashboard. Audit your company for agent opportunities. Build a browser game as a lead magnet.
These are good prompts. Genuinely useful. They are also pure reasoning and research, with no computer use anywhere in them. They ran fine on the previous generation, and they run fine on Claude. Nothing about them requires Astra, and paying $100 a month to run them would be money thrown away. Keep the prompts. Use whatever you already have.
Worth knowing: the game building one is close to a proper recommendation. Games and interactive builds are the most heavily demonstrated Astra category by a wide margin, with community indexes listing 51 builds and 18 with live playable links. What nobody has demonstrated is the lead capture and the funnel behind it. Treat the game as solved and the marketing as your job.
The two that are real
Be my browser operator and the web half of the whole QA team. These are the ones that use what Astra is for, and both have working demonstrations from named people with published artifacts. They are the subject of the next section.
Five Things People Have Actually Proved
Every item below has a named operator and something you can inspect: a screenshot, a video, a live link, or a published result. Sorted by how likely it is to pay for the subscription.
1. Run QA on your own site, funnel, or checkout
This is the strongest recommendation on this page and it is not close. Claire Vo, who runs the product tool ChatPRD, ran Astra through her live application for one hour and forty five minutes. It sent messages, refreshed pages, read the browser console, and surfaced race conditions she had missed. She then fixed them in the same session. Her own summary was direct: if you are not using browser use for QA, please do.
Two things make this the right first experiment. It is genuinely new capability, and the failure mode is harmless. If the agent gets confused on your own staging site, nothing breaks, no money moves, and nobody sues you.
Open [YOUR STAGING OR LIVE URL]. Test every interactive element: navigation links, the contact form with both valid and invalid data, and the checkout flow end to end. Read the browser console after each step. Screenshot anything broken or confusing, and tell me exactly what you did to trigger it.
Start with your checkout or booking flow rather than your whole site. It is where broken things cost you money directly, and it is small enough that the agent can finish inside one run.
2. Prep an edit in Final Cut or Premiere
Two independent demonstrations, which is rare in this material. Video editor Ben Davis had Astra set up a freshly recorded video in Final Cut Pro. It created asset folders, built compound clips, assembled a multicam setup, and applied colour grading, working out the specifics on its own. The editor watching called it more than he expected.
Set up the vid I just recorded in Final Cut, import clips, color grading, sync clips, etc., to get ready to edit.
Separately, Dan Shipper's team at Every had Astra edit a review video unattended for five hours and shipped the result. Their note is the one worth repeating: work like this via computer use was previously impossible.
If you produce video, this is the entry most likely to pay for your subscription. Not the creative edit, the two hours of setup that happens before it.
3. Carry work across two design tools
Claire Vo again. She had Astra inspect an existing node graph in the design tool Flora, import new podcast photos, build correct node graphs at the right aspect ratios, and then assemble the finished thumbnail in Figma. Screenshots of both stages published.
The transferable idea is not thumbnails. It is that Astra can carry work across two separate browser tools with no integration existing between them. Product listing creatives, ad variants, and social cutdowns all have that same shape. If your workflow currently involves exporting from one tool and manually rebuilding in another, that is the candidate.
4. Turn scattered feedback into a working dashboard
This is the proven version of the "one person company dashboard" idea from the viral list. Claire Vo pulled Intercom, GitHub, and Linear into a system that produces an auto-generated wiki, a ranked top three priorities, and trend graphs. It worked in three prompts, after six months of failing at the same task with earlier models.
For a small business, the translation is your support inbox, your review platforms, and your call notes into one weekly read. Feed it exports rather than connecting live accounts, for reasons covered in Section 07.
5. Clean up a library or a codebase
The clearest published prompt in the whole corpus, from the developer known as Theo, with a screenshot of more than forty performance pull requests generated overnight.
Go through my entire codebase hunting for slop, useless tests, unnecessary function wrappers, dead code. Remove them, run the tests, and verify nothing broke.
This one runs in Codex, which Plus users do have access to. The shape generalises well past code: any large collection with accumulated junk and some way to verify you did not break anything.
One thing to not use it for. If you sell writing, Astra is not the upgrade. Three independent negative results: an editorial voice ranking placed it eleventh, below OpenAI's own previous model. A test asking it to write in a specific person's style was caught immediately by an AI detector. And a finance writer called the output obvious and inflated, with instantly recognisable machine prose. The jump here is in hands, not in voice.
How to Prompt It So It Finishes
This is the most useful thing published in Astra's first week, and it got almost no attention next to the viral prompt lists.
Astra's characteristic failure is not a wrong answer. It is stopping early, over-testing, or reporting a job done that it did not actually do. One developer filed an issue documenting Astra reporting that "compile check and diff check passed" twenty two seconds into work that needed minutes of tool calls. It had fixed one bug out of four and claimed all four were resolved.
That is the worst possible failure for a small business owner. A crash you notice. A false completion report you do not. These six patterns target it directly.
Tell it to act instead of asking
Its default is to stop and check with you. On a long task that means a dozen interruptions and a burned allowance.
Infer my intent and task scope from my prompt and our prior conversation. Bias towards action, make reasonable assumptions, and only stop to ask me if a choice is genuinely irreversible.
Right-size the verification
Left alone it will write and run broad test suites nobody asked for, spending your quota on work that never ships.
For this task, use the smallest verification that matches the risk. Do not add or run broad test suites unless I ask.
Define what done actually means
This is the one that catches false completion. Make it show you what it saw, not what it meant to do.
Treat this as complete only when you have implemented the requested change and run or inspected the result yourself. Show me the actual output you observed, not a summary of what you intended to do.
Give it permission up front
Stating that an environment is safe removes a whole category of stopping behaviour.
The local tests use disposable fixtures and have no production access. Run them, fix any failures, and continue without asking me.
Default to medium reasoning, not high
Reasoning tokens bill as output, at $50 per million on the API. High effort on a task that did not need it is money spent on thinking you will never read. Start at medium and raise it only when medium visibly fails.
Never say "make it better" or "test everything"
Both are named anti-patterns by practitioners with real access. Vague scope is what produces the sprawling, over-engineered changes several early testers reported. One team documented Astra making sweeping changes instead of minimal fixes, burning tokens on unnecessary research, and abandoning tasks earlier than competing models.
Stack the first four. They are not alternatives. Put all four instructions at the top of any long-running task and you remove most of the reasons a run dies halfway. This single habit is worth more than any prompt on the viral list.
Your First Week With Astra
If you decide to run the experiment, here is a sequence that gets you a real answer in about a week without wasting your allowance.
- Day 1: Find itOpen ChatGPT and check whether Astra appears in your normal chat model picker. If you are on Plus and it does not, look in ChatGPT Work and in Codex. Do not assume it is missing until you have checked all three surfaces.
- Day 1: Pick one boring jobChoose one thing you or somebody on your team does by hand, in a browser or a desktop app, that takes more than thirty minutes and repeats. Not your most important job. Your most tedious one.
- Day 2: Run the QA test firstPoint it at your own checkout or booking flow with the prompt in Section 04. This calibrates you on what the agent is actually like to work with, at zero risk.
- Day 3: Run your boring job with the four patternsTake the job from Day 1 and stack the four prompting instructions from Section 05 on top of it. Watch the first run. Note where it stopped and why.
- Day 4: Run it again with the fixesAlmost nothing works first time. The second run, with the stopping points addressed, is the run that tells you whether this is real for your business.
- Day 5: Do the arithmeticHow long did the job take by hand? How long did it take including your supervision? Multiply by how often it happens. If the saving does not clear $100 a month, cancel Pro and revisit in three months.
Supervise every run. Nobody has published a real world figure for how long an unattended run lasts before it needs a human, and any confident number you see is invented. What is known is that roughly one task in four fails on the benchmark, and that failure sometimes looks like a confident report of success. Watch the first several runs of anything before you consider leaving it alone.
The Money and the Guardrails
The pricing cliff nobody mentions
API pricing is $10 per million input tokens and $50 per million output. But requests above 272,000 input tokens are repriced at two times input and one and a half times output, for the entire request, not just the amount over the line. That threshold sits at roughly a quarter of the advertised context window.
| Request size | Cost | What happened |
|---|---|---|
| 250K in, 20K out | $3.50 | Under the threshold. |
| 300K in, 20K out | $7.50 | Twenty percent more input. One hundred and fourteen percent more cost. |
| 1M in, 50K out | $23.75 | Against $12.50 at the headline rates. |
Browsing agents accumulate context faster than anything else you will run. Every screenshot, every page of markup, every retry. Crossing 272,000 tokens on a long browsing run is the normal outcome, not the edge case. Two other costs surprise people: reasoning tokens bill as output, and you pay for runs that get killed partway. At a one in four failure rate, budget for paying twice.
Runs get stopped, and the API cannot resume
OpenAI says its safety checks "can sometimes slow, pause, or stop legitimate work." Inside ChatGPT you get asked to review, and you can continue. Via the API the task simply stops, with no way to pick it back up. Developers currently cannot tell a safety stop from an ordinary timeout.
The three ingredients rule
An agent becomes genuinely dangerous when it has all three of: access to your private data, exposure to untrusted content from the web, and the ability to send things outward. Any two are usually fine. All three, and a poisoned web page or a malicious calendar invite can instruct your agent to send your data somewhere you did not choose.
This is not theoretical. Documented incidents from 2026 include an assistant that exfiltrated data through poisoned log entries, an agent that ignored stop commands and deleted a user's emails, and a weaponised calendar invite that planted instructions which fired later when the user asked for a summary of their day. One security team found ten verified attack payloads sitting on unrelated live websites, including instructions for recursive file deletion and $5,000 payment transfers.
Astra is the most injection-resistant model OpenAI has shipped, and that is still not enough. Its attack success rate on indirect prompt injection is 8.5 percent, roughly one in twelve. Against an attacker adjusting tactics over several turns, its defence rate is about 67 percent.
Do not connect any of these: business banking or payment rails with write access, your primary email with permission to send, any password manager or credential store, production systems or customer databases with a delete path, accounts whose loss would end your business such as Stripe or your domain registrar, and client data you are contractually obliged to protect. Your existing NDAs almost certainly did not contemplate this.
If an agent makes a deal, you probably own it
Under existing electronic contracting law, a contract can form through the interaction of electronic agents even when no person reviewed it. The legal analysis is blunt: a principal may not be able to void a contract by explaining afterward that the computer was not supposed to make it. An agent using your email and your name looks fully authorised to the other side, whatever your internal limits were.
Applied to that bill negotiator prompt: if your agent accepts a twenty four month lock-in to save $15 a month, that is likely your contract. Anything touching money or legal consequence needs a human read before it goes out. OpenAI's own disclaimer says ChatGPT is not a fiduciary, adviser, or law firm, and that you are responsible for your decisions.
Put September 15, 2026 in your calendar. Cloudflare sets new defaults on that date that block Training and Agent category bots on ad-carrying pages, for domains newly onboarding to their network. Existing customers keep their settings and are being notified in advance. It is not a blanket switch, but it is the direction of travel, and some of the browser agent demos circulating right now will quietly stop working on certain sites.
Astra or Claude
Both companies publish their own benchmarks, so read every number as a claim rather than a fact. Where a vendor publishes a result they lost on, that carries more weight than one they won.
| Job | Pick | Why |
|---|---|---|
| Browser and desktop automation | Astra | Better scores on the same tests, and roughly 47 percent less time per task. The gap is real and demonstrated. |
| Research and hard reasoning | Claude, narrowly | 65.0 against 57.2 on the hardest reasoning benchmark, published by OpenAI themselves. |
| Long documents | Claude | Both offer around a million tokens of context. Astra reprices the whole request past 272K. Anthropic applies no length surcharge. |
| Coding | A genuine tie | 58.18 against 57.88 on the public leaderboard. Pick on workflow, not score. |
| Building reusable workflows | Claude | OpenAI replaced Custom GPTs with Workspace Agents, which are Business tier and above. A solo operator on Plus has no reusable agent layer at all. |
| Frontier model on a cheap plan | OpenAI | ChatGPT Plus at $20 includes Astra. Claude Pro at $20 does not include Fable 5.1, which needs a Max plan, a Premium seat, or credits. |
The number everyone is quoting wrong
You will see Astra credited with 99.9 percent on a test called ARC-AGI-3, often described as human-level. That score came from a harness that preserves the model's reasoning state between requests and lets it reuse prior work. On the provider-neutral standard harness, the same model scores 62.7 percent. The organisation that runs the benchmark says these results should be understood as the combined performance of the model and its tools. The two figures were also run at different effort settings, so even the gap is not a clean comparison.
Running both
Two $20 subscriptions is a legitimate setup and costs less than one higher tier seat. The split that makes sense is Astra as the hands, doing collection, form filling, and repetitive interface work, and Claude as the reasoning and build layer, doing synthesis, long documents, and the reusable Skills that encode a process you repeat.
What does not exist is any product that hands a task from one to the other automatically. You are the connection between them, or files are.
What To Do Monday Morning
The most useful thing to take from all of this is not a prompt. It is a habit.
When the next model launches, and there will be one, a list of impressive-sounding prompts will circulate within forty eight hours. Before you spend a weekend on it, ask three questions. Has anybody published a result, or only the prompt? Does the prompt need a capability the vendor documents as existing? And does it need the new model at all, or would the one you already pay for handle it?
Those three questions would have caught seven of the nine prompts in this article.
As for Astra itself, the honest position is this. It is a real advance in one narrow area, that area is operating software, and that is worth money to a lot of small businesses who currently pay a person to click through the same screens every week. It is not close to a system you point at your accounts and walk away from, and anybody telling you otherwise has not run it.
Go run the QA test on your own checkout. It costs one afternoon and it will tell you more about where this technology actually is than another month of reading threads about it.
A note on dates. Everything here was checked against primary sources on September 9, 2026, six days after Astra launched. Some of it will change, and some of it will change quickly. Where a figure was a rumour we said so, rather than rounding it into a fact.