Experiments

Comparing 7 Ways to Build 5 Slides: The Most Expensive Scored Lowest

12 min read#experiment#ai#slides#automation

Who this is forAnyone considering making presentation slides with AI, and anyone deciding whether to keep maintaining their own slide generation code.

TL;DR: I made the same five slides seven different ways and measured cost, time, and spec compliance. Comparing only the amounts actually billed, there was a 6x gap between $0.67 and $4.07, and the more expensive path followed the spec the least. Adding subscription cost divided by usage as a converted figure widens the gap to 45x. There is also one difference my scoring rubric cannot catch, so I added another axis to measure.

The number of tools for making presentation slides has exploded. For several months I have been building and using a slide generator with python-pptx. Looking at the tools coming out now, I keep wondering whether I should keep refining this code.

So I measured it. I made the same five slides seven ways and got numbers for how much each cost, how long it took, and how well it followed the spec I set.

Contents

  1. Experiment design
  2. Results: the most expensive path did the worst
  3. Image generation follows the color codes
  4. What my scoring rubric missed
  5. Where Canva MCP broke down
  6. Building costs and fixing costs are different
  7. Summary

Experiment design

For the comparison to hold, all seven paths need exactly the same input. I pinned the 35 slide text strings in one file and fixed the prompt in one file. The only thing that differs between paths is the final paragraph, which says “make it this way in this tool.”

PSubject
Five slides from "How AI and Agents Differ" (cover, cycle diagram, comparison table, big number, closing), 35 text strings
IWhat changed
Six production paths: codex built-in image_gen, Gemini nano-banana-pro, Canva MCP, Genspark standard, Genspark Ultra, Claude Design
CBaseline
Self-built python-pptx generator, code I refined over several months. Whether there is a reason to keep maintaining it is the question of this experiment
OMeasured
Cost (USD, with accounting-basis notes), time in seconds (from submission until spec met), string accuracy (checked against 35 strings), background color error (RGB distance), share of off-spec chromatic color, shape fidelity to the requested form (visual judgment, an axis added after seeing the results)
TPeriod and runs
As of September 9, 2026, second round with color codes included. One main run per path (retry cap of 3, 30-minute cap per path). Fix runs were done once each on four paths only
Not controlled
The 8-color palette is my generator's default, so the control group gets color points for free. The several months spent on the generator are not counted in production time. Run order was fixed, so I became more familiar with the tools as the runs went on. Only Claude Design was operated by hand in the app, so its time is not measured on the same ruler. Genspark's export is paid, so I scored it by capturing the screen
The text is exactly as in spec.json. The prompt is identical character-for-character across common.md (sha256 87f83347…), with the 8-color palette, Noto Sans KR, and 16:9 fixed. All seven paths read one prompt file.

I record the uncontrolled variables openly because that is where the conclusion can change. Why this distinction matters is covered in more depth in why A/B testing was born, which discusses confounding variables and randomization.

The five slides are built around the topic “How AI and Agents Differ.” There is one cover slide, one cycle diagram slide, one three-column, five-row comparison table slide, one big number slide, and one closing slide. I deliberately mixed tables, diagrams, and big numbers because that is where each tool’s inability to build a particular form shows up.

Scoring was done mechanically on two things. The first is string accuracy: I extract text from the output and compare it against the original 35 strings. The second is color compliance: I pull the most frequent colors from the PNG pixels, measure their distance from the spec background color, and count what percentage of the screen is chromatic color outside the 8-color spec palette.

Results: the most expensive path did the worst

Path Cost Time Background color error Off-spec color Shape fidelity
Self-built python-pptx $3.77 actual charge 287 s 0.0 0.6% or less Partial
codex built-in image_gen $0.65 converted 200 s 2.2 or less 1.4% or less Met
Gemini nano-banana-pro $0.67 metered 150 s 7.0 or less 0.5% or less Met
Canva MCP $4.07 actual charge 390 s 22.7 19% or more Not met
Genspark Standard $0.09 converted 180 s 1.0 0.9% or less Not met
Genspark Ultra $0.16 converted 500 s 1.0 0.6% or less Met
Claude Design Not measured 348 s 0.0 0.9% or less Met

I left string accuracy out of the table because all seven paths got a perfect score. Garbled Korean text or changed wording never happened once. What did differ was whether the spec was followed.

Only Genspark is billed in credits. The account I actually ran was on the Free plan, so actual cash spent was $0. Counting only credit deductions, the standard tier used 35 credits for five slides, and the Ultra tier used 65 credits for five slides. To convert this to dollars, I need the paid plan’s unit price. The official pricing page lists Plus at $24.99 per month for 10,000 credits per month, which works out to $0.0025 per credit. So 35 credits is $0.09 and 65 credits is $0.16. Paying annually brings the monthly price to $19.99, which lowers it further.

Seven-path comparison sheet. Rows are paths, and the columns are three slides: the cycle diagram, the comparison table, and the big number. Six rows share a light gray background with a blue accent color, while only the Canva row stands out with a dark navy border and a cream card.
I placed the three slides where capabilities diverge (cycle diagram, comparison table, big number) side by side across the seven paths. Click the image to view it at full size.

Image generation follows the color codes

This is the biggest prediction I got wrong in this experiment. I expected that giving an AI image generator a color code like #2563EB would produce a roughly similar blue. In practice, it drew the table header background fill and the blue underline beneath it exactly to spec.

Slide made with AI image generation. Under the title 'Agents work in a loop,' three boxes labeled observe, judge, and run a tool are arranged in a triangle, connected by curved arrows in a cycle, with a label in the center reading 'repeat until the goal is reached.'
Cycle diagram drawn by codex built-in image_gen. Even the return arrow is accurate.

I want to state the path name precisely. Calling it “GPT image generation” would be wrong. This path runs codex CLI 0.153.4 with the gpt-6-astra model, and the picture was drawn by codex’s built-in image_gen tool. gpt-6-astra is the model that receives instructions and calls the tool, not the model that draws the image. The logs do not record which image model the built-in tool uses behind the scenes, so I did not record an image model name. The other path has a clear name: I called the Gemini API directly and specified nano-banana-pro-preview.

This cycle diagram is also a shape my own code could not build. The diagram types registered in my generator are only straight arrows, so drawing a loop required extra helper code. Owned code has to be modified every time it hits a form outside its vocabulary, and that is the maintenance cost.

Here is the result that hurt most for my generator. Only two paths got a background color error of 0.0: my own code and Claude Design. On the palette, Claude Design actually used all eight colors, while my code used seven. It matched the colors as precisely as my code did, and it also built a loop my code could not. The advantage that months of refinement had bought my code was accuracy, and that advantage did not hold here.

Side-by-side comparison of cycle diagram slides from the two paths that got a background color error of 0.0. On the left, self-built python-pptx has three boxes in a horizontal row with a dotted line returning below. On the right, Claude Design has three boxes arranged in a triangle with three arrows forming a loop.
Same background color error of 0.0. Left is my generator, right is Claude Design.

The two paths that drew images, codex and Gemini, get different notes. Both failed to meet the requested resolution. Codex reported this itself:

Resolution requirement not met. I requested 1920×1080, but the tool returned 1672×941. Code rendering was not used.

Also, because the output is an image, I cannot mechanically extract and verify the text. That means I cannot automate proofreading for typos.

What my scoring rubric missed

Judged by the rubric alone, the two Genspark tiers produced the same result. So I concluded that what the more expensive tier bought was just more retries, not a better design. That is, until I put the outputs side by side.

When I ran Genspark in the standard and Ultra tiers, all the original dependent variables were effectively tied. Background color error was the same at 1.0, palette hit rate was the same, and required elements were 5/5 for both. But when I looked at the outputs side by side, they differed.

Side-by-side comparison of cycle diagram slides made by the Genspark standard and Ultra tiers. On the left, the standard tier has three boxes in a horizontal row with a curve without an arrowhead passing below. On the right, the Ultra tier has three boxes arranged in a triangle, with three arrows forming an actual loop.
Left is standard, right is Ultra. The spec called for a "looping cycle," but the left is a single row.

The problem is that the required-elements check only counted whether something existed. It asks “are there three boxes?” but not “are the boxes arranged in a loop, with an arrow returning from the last box to the first?” So the version that lined the boxes up in a row and drew a curve without an arrowhead also received a perfect score.

So I added another dependent variable: shape fidelity. For slides where the spec specifies a form, I check three things.

Check Condition for “not met”
Three boxes exist as shapes Replaced by a text list
Arrangement is a loop (circle or triangle) Straight single-row arrangement
Arrow from the last box back to the first Line without an arrowhead, or none

Re-scoring with this lens changed the ranking. codex, Gemini, Ultra, and Claude Design are met. Self-built code is partial. Genspark Standard and Canva are not met.

The more expensive tier did not buy only more retries. Retry differences did exist. The standard tier’s comparison table overflowed the bottom of the screen by 189px, so its self-check loop ran twice and it fixed the overflow itself. Ultra produced all five slides with no overflow on the first attempt. On top of that, the diagram’s form matched the spec better.

If 65 holds, the difference between standard and Ultra is $0.09 versus $0.16, a gap of $0.07. At that level, I would skip standard and go with Ultra. The standard tier fails to match the form on slides that include a diagram, and the cost of rerunning to fix that exceeds $0.07. The trade-off is that Ultra is 2.8x slower, so for urgent work I would use standard.

Where Canva MCP broke down

It was the most expensive path and took the longest, yet it produced the worst results. The cause was the structure of its editing API.

Comparison table slide made with Canva. Inside a dark navy border sits a cream card, with five lines of pipe-delimited text in place of a table. There are no grid lines and no footer.
A table was requested, but pipe-delimited text came out.

What the Canva editing API provides is text replacement, element deletion, position and size changes, styling, and image filling. There is no feature to add new elements. So you can only reuse elements that already exist in the template produced by automatic generation. If the template has no table or shape, there is no way to build that slide. There is also no feature to change the background color, so the template’s dark navy was carried over as-is.

Building costs and fixing costs are different

What actually eats time in practice is not the first build but the fixes. So I asked for one table row to be added to the finished five slides. This fix run was done only on four paths. I did not measure fixing costs for the two Genspark tiers or for Claude Design.

Path Fix cost Time Result
Gemini nano-banana-pro $0.134 24 s Row added. Footer contaminated
codex built-in image_gen $0.83 150 s Row added. No effect on other slides
Self-built python-pptx $1.21 68 s One line of data fixed, then rebuilt
Canva MCP $1.75 110 s Row added. Since it is text rather than a table, only one line was added

My prediction was off here too. I expected that regenerating an image would make its tone clash with the other slides, but the background color shift was within 0.2. That is not distinguishable by eye.

Instead, a different problem appeared. Gemini added the row correctly, but the design spec instructions leaked straight into the footer.

Lower part of a comparison table slide regenerated by Gemini. The bottom-right footer begins with a font specification and a color code, followed by a copyright line.
Noto Sans KR, 20px, #1E293B. came out attached. The instructions leaked into the output.

Image generation has no boundary between instructions and content. A spec written into the prompt can appear as text inside the picture, and there is no mechanical way to catch it.

Summary

This post is one entry in an experiment log that runs every two weeks. In the next installment, I plan to measure other tasks the same way. You can see just this series on the Experiment page.

Frequently asked questions

What was the billed cost gap between the cheapest and most expensive paths?
Counting only amounts actually billed, the gap was about 6x, from $0.67 for Gemini nano-banana-pro to $4.07 for Canva MCP.
Did the Canva MCP path follow the slide spec?
No. Canva MCP had a background color error of 22.7, 19% or more off-spec chromatic color, and did not meet the shape fidelity requirement.

Want the full system? The Claude Code & Codex Skills guidebook collects the skills and subagents behind this blog, from $19.