AI PowerPoint Agent Benchmark: 12 Tests That Matter

A reproducible benchmark for AI PowerPoint agents covering native objects, existing-deck edits, charts, sources, recovery, visual QA, and delivery.

Bob · Former McKinsey and Deloitte consultant with 6 years of experienceJuly 19, 20268 min read

Pricing and feature information was accurate at the time of publication. Competitor products change frequently — verify current details on each provider's website.

Most AI presentation comparisons score the easiest thing to see: the first screenshot.

That misses what makes a PowerPoint agent useful. A professional agent must understand an existing deck, create native objects, apply revisions without destroying unrelated work, recover from imperfect steps, preserve sources, and deliver a file that a human can keep editing.

This benchmark tests those outcomes. It does not publish vendor scores yet. A credible comparison requires every product to be run on the same date, plan, input deck, prompts, and evidence standard. The methodology comes first.

AI PowerPoint benchmark across creation, revision, verification, and delivery

Benchmark Principles#

Test the artifact, not the marketing page#

Record what exists in the final .pptx: object types, chart data, source markers, slide structure, formatting preservation, and manual editability.

Test continuation, not one prompt#

An agent should perform a sequence of dependent actions and recover when one step is imperfect. A single prompt-to-deck task mainly tests generation.

Use the same evidence bundle#

For each product, save:

  • Product, plan, version, date, and environment
  • Original input deck
  • Source files and structured data
  • Exact prompts and follow-ups
  • Tool or activity transcript when available
  • Final PowerPoint file
  • Rendered images of every changed slide
  • Native-object inspection notes
  • Timing and human interventions
  • Pass/fail evidence for every criterion

Separate capability from quality#

A task can be supported but executed poorly. Score both:

  • Capability: Could the product attempt the task through a relevant workflow?
  • Outcome: Did the delivered artifact satisfy the rubric?

The Test Pack#

Use a 12-slide consulting-style deck containing:

  • An executive summary
  • Two section dividers
  • A native bar chart with source note
  • A formatted table
  • A slide with icons and grouped shapes
  • A market-sizing slide with three explicit assumptions
  • A recommendation slide
  • A small appendix

Include three source files:

  1. A two-page market brief with citations
  2. A CSV containing updated quarterly values
  3. A one-page review memo with five requested changes

The deck should contain one deliberate defect, such as a clipped label, so visual review can be tested.

The 12 Tests#

Test 1: Deck understanding#

Prompt: “Summarize the deck's argument in five bullets. Identify the two weakest evidence links and the slide numbers involved. Do not change the presentation.”

Pass criteria: Correct storyline, correct slide references, no mutation, and uncertainty stated where evidence is weak.

Why it matters: An agent cannot safely edit state it does not understand.

Test 2: New-slide creation#

Prompt: “Using the market brief, add one slide that argues for entering through a local partnership. Include three supporting reasons and visible source references. Match the surrounding deck.”

Pass criteria: Slide exists in the right section, makes the requested claim, uses grounded evidence, and fits the deck's visual language.

Artifact check: Text and design elements are individually editable.

Test 3: Native chart creation#

Prompt: “Create a native waterfall chart from the supplied CSV showing the bridge from Q3 to Q4. The message is that pricing and volume offset FX and returns.”

Pass criteria: Values and labels are correct, totals reconcile, chart is native and editable, and the action title reflects the intended message.

Failure condition: A screenshot of a chart is not a native-chart pass.

Test 4: Existing-chart update#

Prompt: “Replace the current chart data with the revised quarter in the CSV. Preserve the chart's formatting and update any dependent claim.”

Pass criteria: Data changes without flattening, unrelated styling remains stable, and dependent text is updated accurately.

Test 5: Targeted existing-slide edit#

Prompt: “Make the recommendation on slide 9 more direct. Preserve its evidence, layout, and all other slides.”

Pass criteria: Only the necessary objects change; the revised title is genuinely more direct; no collateral edits appear.

Test 6: Cross-slide consistency#

Prompt: “Replace ‘Southeast Asia expansion’ with ‘Southeast Asia entry’ wherever it refers to the recommendation, but do not change quotations or source titles.”

Pass criteria: Intended occurrences change, protected text remains unchanged, and the agent reports any ambiguous cases.

Test 7: Source replacement#

Prompt: “Replace the old market-growth source with the supplied brief. Update the affected citation and list any claim that is no longer supported.”

Pass criteria: Source reference changes, unsupported claims are identified rather than silently preserved, and unrelated citations remain intact.

Test 8: Review memo execution#

Prompt: “Apply the five changes in the review memo. Preserve completed work and report each change with its slide number.”

Pass criteria: All five items are resolved or explicitly blocked, changes map to the correct slides, and the final report matches the artifact.

Test 9: Failure recovery#

Include one request the product cannot safely perform, such as a host-specific animation or unavailable data link.

Pass criteria: The agent preserves successful work, does not fabricate completion, explains the blocker, and proposes a bounded alternative when appropriate.

Failure condition: Repeating the same action without progress or claiming success without evidence.

Test 10: Visual verification#

Prompt: “Inspect every changed slide for clipping, overlap, weak hierarchy, and visual inconsistency. Fix the deliberate label defect and report what you verified.”

Pass criteria: The seeded defect is found and corrected, changed slides remain legible, and reported checks correspond to rendered evidence.

Test 11: Manual handoff#

Open the output in PowerPoint and perform five manual edits:

  1. Change the new slide title
  2. Edit one chart value
  3. Move one icon
  4. Restyle one table cell
  5. Apply a theme color

Pass criteria: All edits are possible through normal PowerPoint interactions without rebuilding or ungrouping a flattened slide.

Test 12: Delivery proof#

Prompt: “Deliver the completed deck and summarize only verified changes and unresolved blockers.”

Pass criteria: The received file is the same artifact that was inspected; the change summary is accurate; no blocking defect is omitted; no unsupported completion claim appears.

Generate consulting slides with AI

Describe what you need. AI generates structured, polished slides — charts and visuals included.

Scoring Rubric#

Score each test from 0 to 3.

ScoreMeaning
0Unsupported, unsafe, or no meaningful result
1Partial result with major errors or manual reconstruction
2Outcome mostly correct with bounded cleanup required
3Outcome correct, preserved, verified, and ready for normal manual continuation

Weight the dimensions for a professional PowerPoint workflow:

DimensionWeight
Outcome accuracy20%
Existing-work preservation15%
Native artifact quality15%
Data and source integrity15%
Recovery and continuation10%
Visual quality10%
Manual handoff10%
Delivery evidence5%

Publish the raw task scores alongside any total. A weighted number without evidence hides trade-offs.

Record Human Effort Separately#

Measure three kinds of time:

  • Agent elapsed time: from prompt to delivered result
  • User coordination time: clarifications, retries, and manual tool operation required to make the agent continue
  • Cleanup time: manual work needed before the deck is usable

A generator that finishes in two minutes but requires forty minutes of cleanup is not faster than an agent that takes ten minutes and delivers native, verified work.

Also record the number and type of interventions. One strategic correction is different from six retries caused by lost state.

How to Compare PowerPoint, Web, and HTML Tools Fairly#

Do not give every artifact the same native-object score.

Native PowerPoint#

Expect real .pptx objects, Office compatibility, and ordinary manual editing.

Proprietary web presentation#

Evaluate editing inside the product, collaboration, hosted delivery, and export fidelity. Do not pretend the hosted source is equivalent to a native PowerPoint file.

Local HTML presentation#

Evaluate source ownership, local preview, browser rendering, version control, visual editing, and the ability for an agent to continue against the files. Deckary Canvas, Reveal.js, Slidev, and Marp belong in this artifact family even though their editing models differ.

The right winner depends on the required delivery format.

What Not to Publish#

Avoid these common comparison failures:

  • “We tested” claims without saved prompts and output files
  • Vendor scores collected on different dates or paid plans
  • Design scores based on one cherry-picked slide
  • Native PowerPoint claims based on file extension alone
  • Speed claims that exclude user coordination and cleanup
  • Agent claims based only on multi-turn chat
  • Overall rankings that hide data, source, or editability failures

If a result cannot be reproduced, call it an observation rather than a benchmark.

Deckary's Quality Bar#

Deckary's own PowerPoint agent should be held to the same method. The product thesis—durable state, explicit tools, native artifacts, sources, visual review, and delivery evidence—is valuable only if it survives a real deck.

Publishing the method before publishing scores creates a useful constraint: future comparisons must show the input, output, and evidence rather than relying on category language.

The benchmark also makes model-provider changes safer. A provider-independent harness should be able to run the same test pack before and after a model or adapter change and show whether outcomes improved.

Generate consulting slides with AI

Describe what you need. AI generates structured, polished slides — charts and visuals included.