Aside Browser: Should You Connect Claude or GPT? Accuracy, Speed, and Cost Test
Who this is forReaders who connect a Claude or ChatGPT subscription to the Aside browser to handle web tasks and can't decide which model to use.
TL;DR: I connected four models to the Aside browser one at a time and gave each the same three tasks nine times. Claude Sonnet 5.5, Claude Opus 5.5, and GPT-5.6 Sol scored perfectly all nine times. The cheapest, GPT-6 Luna, missed twice on the shopping comparison, where several conditions had to be checked. Luna got every search and form-filling run right. If you have a Claude subscription, Sonnet was the safest pick in this test; with only a ChatGPT subscription, Sol was.
Aside is an AI agent browser that lets you connect the Claude or ChatGPT subscription you already have. When I opened the settings to hand web tasks to an AI agent, a question came up right away. If both are connected, which model should get the job?
The official benchmark that Aside published includes results for only one GPT configuration, with no table comparing models (checked October 6, 2026). So I measured it myself.
Contents
- Models to connect to Aside and per-task settings
- Experiment design
- Results: three models scored perfectly all nine times
- Which one to choose in each situation
- Two ways Luna got it wrong
- Form filling: all four models succeeded
- Cache write costs for the Claude connection
- Limitations of this experiment
- Summary
1. Models to connect to Aside and per-task settings
Connect your subscriptions on the Models screen in Aside’s settings. I connected both my Claude subscription and my ChatGPT subscription.
On the same screen, under Task models, you can assign a different model to each type of task.
| Slot | Tasks it handles | My setting |
|---|---|---|
| Default | Default model for new chats | GPT-6 Luna |
| Fast | Light, short side tasks | GPT-6 Luna |
| Standard | Background tasks, memory cleanup | GPT-5.6 Te (cut off on screen) |
| Deep | Hard reasoning, planning, judging | GPT-5.6 Sol |
| Visual | Images, screenshots, page interpretation | GPT-5.6 Sol |
| Image generation | Image generation | Not used in this experiment |
The steps for connecting a subscription are covered with screenshots in how to use the Aside browser.
2. Experiment design
For the comparison to hold, everything except the model had to stay the same. I fixed the task prompts in one file, and the only thing I changed between runs was the one model connected.
- PPopulation
- Three browser tasks: web search (live web), shopping comparison (mock shop), booking form entry (mock booking form)
- IIntervention
- Four models connected to Aside: GPT-6 Luna, GPT-5.6 Sol (GPT connection), Claude Opus 5.5, and Claude Sonnet 5.5 (Claude connection)
- CComparison
- GPT-6 Luna, the default model in my settings and the cheapest model
- OOutcomes
- Accuracy (grader checked against correct answers), time (seconds, from the first session log entry to the last), cost (USD converted at official API prices), model shown in the log, and recovery after form errors
- TTimeframe and runs
- October 6, 2026. Three main runs per model-task combination, 36 in total (about 15 minutes), plus 4 form-revision runs, for 40 in all
- Not controlled
- Aside's memory may carry over from one run to the next (I rotated the order so no side was favored); background costs such as memory cleanup may fall outside the session log; the amount of reasoning may differ between companies even at the same medium setting
I picked three tasks that anyone does in a browser. The shopping mall and booking form are mock sites that run only on my MacBook, so no real orders or reservations were created.
| Task | What it does | Target | How correctness was judged |
|---|---|---|---|
| Web search | Answer 3 public-information questions and cite sources | Live web | Checked against values on official agency pages |
| Shopping comparison | Pick the product that meets 4 conditions out of 12 earbuds | Mock shop | 3 correct answers calculated from the conditions |
| Booking entry | Enter values into a 2-step booking form and submit | Mock booking form | 10 submitted values saved on the server |
| After form revision | Add one required field to the same form | Mock booking form | 11 submitted values, once per model |
Costs are calculated from the tokens recorded in the logs at official API prices. I connected both companies through subscriptions, so there were no actual charges. Please read every cost below as what it would have cost via the API.
3. Results: three models scored perfectly all nine times
| Connection | Model | Perfect runs | Average score | Average time | Cost per run (API equivalent) |
|---|---|---|---|---|---|
| Claude | Claude Opus 5.5 | 9/9 | 1.000 | 20.9 seconds | $0.318 |
| Claude | Claude Sonnet 5.5 | 9/9 | 1.000 | 16.0 seconds | $0.180 |
| GPT | GPT-5.6 Sol | 9/9 | 1.000 | 33.8 seconds | $0.182 |
| GPT | GPT-6 Luna | 7/9 | 0.924 | 28.0 seconds | $0.005 |
4. Which one to choose in each situation
5. Two ways Luna got it wrong
The shopping comparison had four conditions: a discounted price of KRW 150,000 (South Korean won) or less, active noise cancellation, battery life of 8 hours or more, and a rating of 4.3 or higher. The list shows only prices and ratings. Battery life and noise cancellation appear only when you open each product’s detail page.
The two times Luna got it wrong were wrong in different ways.
- Run 1, boundary judgment: It opened all 12 detail pages but included Airbeat C3, with its rating of 4.29, and left out Harmony B2, whose discounted price was exactly KRW 150,000.
- Run 2, stopped searching: After opening 2 detail pages, it stopped when a page element error occurred. It never opened page 2, so Audiopoint G7 was missing, and it admitted on its own that the answer was incomplete.
- Run 3: Correct. The other three models were correct in all nine runs.
However, I haven’t checked the detail pages of the remaining products or page 2 yet, so this result can’t be confirmed against the full product list.
(Aside session log, GPT-6 Luna shopping comparison run 2, final answer)
6. Form filling: all four models succeeded
I deliberately put one trap in the booking form. If you enter a phone number without hyphens, the form returns a format error.
All four models got this error every time. Each time, they corrected the number on their own to 010-1234-5678 and resubmitted. No run required human intervention.
Your reservation is complete. The first submission returned a phone number format error, so I corrected it to 010-1234-5678 as instructed and resubmitted.
(Aside session log, Claude Sonnet 5.5 booking entry run 1)
I then added one more required allergy field to the form and gave the task again. Even though the screen had changed, all four models submitted all 11 values correctly.
| Model | Submitted values | Time | Cost (API equivalent) |
|---|---|---|---|
| Claude Sonnet 5.5 | 11/11 | 16 seconds | $0.153 |
| Claude Opus 5.5 | 11/11 | 18 seconds | $0.266 |
| GPT-5.6 Sol | 11/11 | 38 seconds | $0.167 |
| GPT-6 Luna | 11/11 | 70 seconds | $0.006 |
Luna gave the correct answer on the changed form too, but it took the longest at 70 seconds. Each model was timed only once, so treat the speed differences as reference only.
7. Cache write costs for the Claude connection
The Claude connection writes at least about 24,000 to 27,000 tokens to a 1-hour cache in each session. In tasks that read a lot, such as web search, that rose to 37,000 to 61,000 tokens. The GPT connection has no cache write cost.
- Across the nine-run average, about $0.26 per run for Opus 5.5 and about $0.14 for Sonnet 5.5 were cache writes (based on the official 1-hour cache write price).
- These tasks were short, taking between 9 and 70 seconds each, so the cache write share was larger than the share spent on the actual work.
- Connecting through a subscription doesn’t bill this money to you. However, I didn’t measure how much of your subscription limit gets used.
8. Limitations of this experiment
- Three runs per task means a small sample. These runs can’t tell whether Luna’s two wrong answers were chance or a trend.
- I measured reasoning effort at medium only.
- For web search, I checked only the numbers in the answer against the correct values. I didn’t grade whether the source URLs were correct.
- The shopping mall and booking form are two local mock sites. They don’t include the pop-ups, logins, or slow responses found on real sites.
- Background task costs such as memory cleanup may fall outside the session log.
- I didn’t measure how much of the subscription limit gets used.
9. Summary
Frequently asked questions
- Which should I connect to the Aside browser, Claude or GPT?
- In my test (3 tasks, 9 runs each), Claude Sonnet 5.5, Claude Opus 5.5, and GPT-5.6 Sol scored perfectly every time, and GPT-6 Luna missed twice on shopping comparison. Sonnet 5.5 was fastest at 16.0 seconds on average. The sample is small, so test your own work too (measured October 6, 2026).
- What model does Aside use if I don't choose one?
- It runs on whatever default model is set at that time. On the same day, I ran on Opus in the morning and on Luna in the afternoon after changing the setting. When comparing models, you must specify the model for every run.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›