Experiments

Aside Browser: Should You Connect Claude or GPT? Accuracy, Speed, and Cost Test

8 min read#aside-browser#ai-agent#claude#chatgpt#model-comparison

Who this is forReaders who connect a Claude or ChatGPT subscription to the Aside browser to handle web tasks and can't decide which model to use.

TL;DR: I connected four models to the Aside browser one at a time and gave each the same three tasks nine times. Claude Sonnet 5.5, Claude Opus 5.5, and GPT-5.6 Sol scored perfectly all nine times. The cheapest, GPT-6 Luna, missed twice on the shopping comparison, where several conditions had to be checked. Luna got every search and form-filling run right. If you have a Claude subscription, Sonnet was the safest pick in this test; with only a ChatGPT subscription, Sol was.

Aside is an AI agent browser that lets you connect the Claude or ChatGPT subscription you already have. When I opened the settings to hand web tasks to an AI agent, a question came up right away. If both are connected, which model should get the job?

The official benchmark that Aside published includes results for only one GPT configuration, with no table comparing models (checked October 6, 2026). So I measured it myself.

Contents

  1. Models to connect to Aside and per-task settings
  2. Experiment design
  3. Results: three models scored perfectly all nine times
  4. Which one to choose in each situation
  5. Two ways Luna got it wrong
  6. Form filling: all four models succeeded
  7. Cache write costs for the Claude connection
  8. Limitations of this experiment
  9. Summary

1. Models to connect to Aside and per-task settings

Connect your subscriptions on the Models screen in Aside’s settings. I connected both my Claude subscription and my ChatGPT subscription.

Models screen in Aside settings. The Providers list has three rows: Aside Free, Claude Subscription, and ChatGPT Subscription, with a Connect button below
Both the Claude and ChatGPT subscriptions are connected in the Providers list.

On the same screen, under Task models, you can assign a different model to each type of task.

Slot Tasks it handles My setting
Default Default model for new chats GPT-6 Luna
Fast Light, short side tasks GPT-6 Luna
Standard Background tasks, memory cleanup GPT-5.6 Te (cut off on screen)
Deep Hard reasoning, planning, judging GPT-5.6 Sol
Visual Images, screenshots, page interpretation GPT-5.6 Sol
Image generation Image generation Not used in this experiment

The steps for connecting a subscription are covered with screenshots in how to use the Aside browser.

2. Experiment design

Illustration of a person seated at a desk with four small helpers standing before them, each holding the same checklist. Beside each helper is a stopwatch and a coin purse of a different size
I gave the same errand to four models and lined up their times and costs side by side.

For the comparison to hold, everything except the model had to stay the same. I fixed the task prompts in one file, and the only thing I changed between runs was the one model connected.

PPopulation
Three browser tasks: web search (live web), shopping comparison (mock shop), booking form entry (mock booking form)
IIntervention
Four models connected to Aside: GPT-6 Luna, GPT-5.6 Sol (GPT connection), Claude Opus 5.5, and Claude Sonnet 5.5 (Claude connection)
CComparison
GPT-6 Luna, the default model in my settings and the cheapest model
OOutcomes
Accuracy (grader checked against correct answers), time (seconds, from the first session log entry to the last), cost (USD converted at official API prices), model shown in the log, and recovery after form errors
TTimeframe and runs
October 6, 2026. Three main runs per model-task combination, 36 in total (about 15 minutes), plus 4 form-revision runs, for 40 in all
Not controlled
Aside's memory may carry over from one run to the next (I rotated the order so no side was favored); background costs such as memory cleanup may fall outside the session log; the amount of reasoning may differ between companies even at the same medium setting
One prompt file per task, reasoning effort set to medium, a new session for each run, and the order of tasks and models rotated between runs. Everything ran one at a time on the same MacBook. The model shown in the log matched the intended model in all 40 runs.

I picked three tasks that anyone does in a browser. The shopping mall and booking form are mock sites that run only on my MacBook, so no real orders or reservations were created.

Task What it does Target How correctness was judged
Web search Answer 3 public-information questions and cite sources Live web Checked against values on official agency pages
Shopping comparison Pick the product that meets 4 conditions out of 12 earbuds Mock shop 3 correct answers calculated from the conditions
Booking entry Enter values into a 2-step booking form and submit Mock booking form 10 submitted values saved on the server
After form revision Add one required field to the same form Mock booking form 11 submitted values, once per model

Costs are calculated from the tokens recorded in the logs at official API prices. I connected both companies through subscriptions, so there were no actual charges. Please read every cost below as what it would have cost via the API.

3. Results: three models scored perfectly all nine times

Connection Model Perfect runs Average score Average time Cost per run (API equivalent)
Claude Claude Opus 5.5 9/9 1.000 20.9 seconds $0.318
Claude Claude Sonnet 5.5 9/9 1.000 16.0 seconds $0.180
GPT GPT-5.6 Sol 9/9 1.000 33.8 seconds $0.182
GPT GPT-6 Luna 7/9 0.924 28.0 seconds $0.005
Cost per run differs by up to 68 times. Horizontal bars starting at 0 show average cost per run, in US dollars, as if paid via API. Opus 5.5 at $0.318, perfect 9 of 9. Sol 5.6 at $0.182, perfect 9 of 9. Sonnet 5.5 at $0.180, perfect 9 of 9. Luna 6 at $0.005, perfect 7 of 9. Only Luna 6 is blue, and its bar is barely visible
Luna's cost per run was $0.0047, about 1/68 of Opus's $0.3182, and it was perfect in 7 of 9 runs.
Sonnet was the fastest. Horizontal bars starting at 0 show average time per run. Sol 5.6 at 33.8 seconds, Luna 6 at 28.0 seconds, Opus 5.5 at 20.9 seconds, Sonnet 5.5 at 16.0 seconds. Only Sonnet 5.5 is blue
The two Claude-connected models were faster than the two GPT-connected models.
Only one cell was not perfect. A grid of models and tasks showing perfect runs. Opus 5.5, Sonnet 5.5, and Sol 5.6 each scored 3/3 on web search, 3/3 on shopping comparison, 3/3 on booking entry, and 1/1 after form revision. Luna 6 scored 1/3 on shopping comparison, shown as a dark gray cell, and was perfect on everything else. Luna once misjudged a boundary value and once stopped before opening all the products
The difference among the four models came from only one cell: shopping comparison.

4. Which one to choose in each situation

Which one to choose in each situation. If you have a ChatGPT subscription and the task has a clear right answer, such as search or form entry, choose Luna 6: perfect on all six search and entry runs, about $0.004 per run. Otherwise, if you have a Claude subscription, choose Sonnet 5.5: perfect in all nine runs, and the fastest at an average of 16 seconds. Otherwise, if you have a ChatGPT subscription, choose Sol 5.6: perfect in all nine runs, at 33.8 seconds on average. If none of these apply, pick a model and measure it yourself. This is a small experiment with three runs per task
I split the choices by type of task and the subscriptions you have. The only evidence is the results of these 40 runs.

5. Two ways Luna got it wrong

The shopping comparison had four conditions: a discounted price of KRW 150,000 (South Korean won) or less, active noise cancellation, battery life of 8 hours or more, and a rating of 4.3 or higher. The list shows only prices and ratings. Battery life and noise cancellation appear only when you open each product’s detail page.

Page 1 of a mock shop's wireless earbuds list. Sorion A1 at KRW 149,000, rating 4.6. Harmony B2 at KRW 150,000, rating 4.4. Airbeat C3 at KRW 119,000, rating 4.29. Nova D4 at KRW 99,000, rating 4.5. Clear E5 at KRW 89,000, rating 4.7. Bitwave F6 at KRW 148,000, rating 4.2. The items are laid out in two rows
Harmony B2's discounted price is exactly KRW 150,000, and Airbeat C3's rating is 4.29.

The two times Luna got it wrong were wrong in different ways.

  • Run 1, boundary judgment: It opened all 12 detail pages but included Airbeat C3, with its rating of 4.29, and left out Harmony B2, whose discounted price was exactly KRW 150,000.
  • Run 2, stopped searching: After opening 2 detail pages, it stopped when a page element error occurred. It never opened page 2, so Audiopoint G7 was missing, and it admitted on its own that the answer was incomplete.
  • Run 3: Correct. The other three models were correct in all nine runs.

However, I haven’t checked the detail pages of the remaining products or page 2 yet, so this result can’t be confirmed against the full product list.

(Aside session log, GPT-6 Luna shopping comparison run 2, final answer)

Detail page for Airbeat C3. Discounted price KRW 119,000, rating 4.29 from 2,051 reviews, battery life 10 hours excluding the case, noise cancellation available
A rating of 4.29 does not meet the condition of 4.3 or higher.

6. Form filling: all four models succeeded

I deliberately put one trap in the booking form. If you enter a phone number without hyphens, the form returns a format error.

Step 2 of a mock restaurant booking form. A red box at the top reads: Please enter the phone number in the format 010-0000-0000. The phone field contains 01012345678
Entering a number without hyphens triggers this error.

All four models got this error every time. Each time, they corrected the number on their own to 010-1234-5678 and resubmitted. No run required human intervention.

Your reservation is complete. The first submission returned a phone number format error, so I corrected it to 010-1234-5678 as instructed and resubmitted.

(Aside session log, Claude Sonnet 5.5 booking entry run 1)

I then added one more required allergy field to the form and gave the task again. Even though the screen had changed, all four models submitted all 11 values correctly.

Model Submitted values Time Cost (API equivalent)
Claude Sonnet 5.5 11/11 16 seconds $0.153
Claude Opus 5.5 11/11 18 seconds $0.266
GPT-5.6 Sol 11/11 38 seconds $0.167
GPT-6 Luna 11/11 70 seconds $0.006

Luna gave the correct answer on the changed form too, but it took the longest at 70 seconds. Each model was timed only once, so treat the speed differences as reference only.

7. Cache write costs for the Claude connection

The Claude connection writes at least about 24,000 to 27,000 tokens to a 1-hour cache in each session. In tasks that read a lot, such as web search, that rose to 37,000 to 61,000 tokens. The GPT connection has no cache write cost.

Most of the Claude cost is cache writes. Horizontal bars of average cost per run, split into cache writes (blue) and everything else (gray). Opus 5.5: $0.318, of which 82% is cache writes. Sonnet 5.5: $0.180, of which 76% is cache writes. Sol 5.6: $0.182, no cache writes. Luna 6: $0.005, no cache writes
Most of the cost per run for the Claude connection came from cache writes.
  • Across the nine-run average, about $0.26 per run for Opus 5.5 and about $0.14 for Sonnet 5.5 were cache writes (based on the official 1-hour cache write price).
  • These tasks were short, taking between 9 and 70 seconds each, so the cache write share was larger than the share spent on the actual work.
  • Connecting through a subscription doesn’t bill this money to you. However, I didn’t measure how much of your subscription limit gets used.

8. Limitations of this experiment

  • Three runs per task means a small sample. These runs can’t tell whether Luna’s two wrong answers were chance or a trend.
  • I measured reasoning effort at medium only.
  • For web search, I checked only the numbers in the answer against the correct values. I didn’t grade whether the source URLs were correct.
  • The shopping mall and booking form are two local mock sites. They don’t include the pop-ups, logins, or slow responses found on real sites.
  • Background task costs such as memory cleanup may fall outside the session log.
  • I didn’t measure how much of the subscription limit gets used.

9. Summary

Frequently asked questions

Which should I connect to the Aside browser, Claude or GPT?
In my test (3 tasks, 9 runs each), Claude Sonnet 5.5, Claude Opus 5.5, and GPT-5.6 Sol scored perfectly every time, and GPT-6 Luna missed twice on shopping comparison. Sonnet 5.5 was fastest at 16.0 seconds on average. The sample is small, so test your own work too (measured October 6, 2026).
What model does Aside use if I don't choose one?
It runs on whatever default model is set at that time. On the same day, I ran on Opus in the morning and on Luna in the afternoon after changing the setting. When comparing models, you must specify the model for every run.