Repro Agent
A bug report tells you what someone saw. The Repro Agent tells you whether it still happens. It takes a recording that was analysed in Bug Report mode, reconstructs the flow from the clicks and navigations that were captured, drives that flow headlessly against an origin you control, and reports whether the failure actually fired.
It is built to be believable rather than encouraging. A verdict is only ever “reproduced” when a real signal was observed — the specific console error, the failing request, the missing element the plan named in advance as the thing to watch for. If it cannot make the bug happen, it says so, and that is a useful answer rather than a failure of the tool.
Before you can run it
- The recording must have been analysed in Bug Report mode. The agent works from the report’s steps and the interaction trail; other modes have neither.
- The recording must carry an interaction event log — the clicks and navigations the flow is reconstructed from. Extension recordings always do. An uploaded video only qualifies when its event log was supplied alongside it through the API (
eventLogUploadId); a plain dashboard upload has no trail to re-drive. Widget reports qualify when their site key captures interactions — see below. - You must have verified the target origin in your workspace. We will not drive a site you have not proven you control.
- Pro and above. Free plans cannot run it.
Running it
Open a bug report and choose Check if it reproduces. It is opt-in per report — nothing is re-driven automatically, because driving a flow means issuing real clicks and real requests against a real environment, and that is your decision to make one report at a time.
Or through the API:
curl -X POST https://api.scenerecap.com/v1/recordings/<id>/repro \
-H "Authorization: Bearer sr_live_…" \
-H "Content-Type: application/json" \
-d '{
"targetOrigin": "https://staging.example.com",
"runs": 3,
"bisect": false,
"counterfactual": false
}'targetOrigin— which environment to drive. Must be a verified origin on your workspace. Point it at staging to check a fix before it ships, or at production to confirm a report from a user.runs— how many times to attempt the flow. More runs distinguish a bug that always happens from one that happens sometimes; the verdict carries the rate.bisect— also compute the minimal repro. See below.counterfactual— also test alternative paths. See below.
The run is asynchronous. The result appears on the recording and is available from the API as the repro pass.
What the verdicts mean
- Reproduced — the flow ran and the failure signal fired. The result carries the rate (“3 of 3”), the signal that was observed, and a screenshot at the failing step.
- Not reproduced — the flow ran to completion and the failure never fired. This is a real answer, not an error: it is what you want to see after a fix, and it is worth taking seriously as evidence that a report is stale or environment-specific. And when you want the fix on camera, a Take re-performs the report’s flow against your fixed build and records it passing — evidence you can paste into the issue before closing it.
- Ambiguous — the flow ran but the evidence does not clearly support either conclusion. Read it as “a human still needs to look”, not as a soft yes.
- Blocked — it could not get far enough to find out: a login it cannot pass, a step whose target no longer exists, a consent wall. The result says what stopped it.
Minimal repro
With bisect enabled, a reproduced bug is re-run with steps removed to find the shortest sequence that still triggers it. A twelve-step report frequently collapses to three, which is the difference between an issue an engineer opens and one they postpone. It only runs when the bug reproduced — there is nothing to minimise otherwise.
Export as a Playwright test
Once a run has finished, Download Playwright test on the recording page saves the flow the agent drove as a .spec.ts file for Playwright. It is a regression test: it re-drives the steps and asserts that the bug’s signal — the failing request, the console error or the on-screen text — does not appear. So it fails while the bug reproduces and passes once it is fixed. Commit it next to your other tests and it guards the fix in your own CI.
- When a minimal repro was found, the test uses only those steps.
- The base URL defaults to the origin the agent drove; set
BASE_URLto point it elsewhere. - Review it before committing. The flow is reconstructed from the recording’s clicks and navigations, typed values were chosen by the AI because SceneRecap never records keystrokes, and a click the recorder could only place by screen position is marked
TODOfor you to give a locator. - The file is generated in your browser from the run you already have. Downloading it uses no agent run.
Counterfactuals
With counterfactual enabled, the agent proposes plausible alternative paths to the same goal and drives those too, answering “does this happen on every route, or only this one?”. Each alternative comes back as avoided or still fails, which often localises the bug before anyone has opened the code. And when you already know which variation you mean, say it: the optional “tell the agent to do it differently” instruction makes the re-drive follow your words — “try it with an empty cart”, “skip the coupon step” — while the pass/fail signal stays the recorded failure, so a “reproduced” still means the real bug fired under your variation.
Widget reports: opt-in
Reports filed through the in-page reporting widget carry console errors and failed network requests by default — but no click and navigation trail, so there is no flow to reconstruct, and the honest thing to do is not to guess at one. Those reports are evidence-only and never labelled confirmed reproducible. If you enable Capture clicks & navigation during recordings on a site key (Settings → Reporting widget), reports from that site carry the trail and are re-drivable like any other; the visitor’s consent line names the diagnostic data. The gate is per report, on the trail actually present — an evidence-only report stays evidence-only even after you flip the switch.
Where it runs
In our infrastructure, in a headless browser, on a clean profile — not on the machine of whoever filed the report. It has none of their extensions, cookies, session, locale, network conditions or device. That is exactly why a “not reproduced” verdict is informative: it narrows the bug toward something specific to that environment rather than proving the report wrong. Anything that depends on being signed in as a particular user, or on a physical device, is a poor fit.
Verified origins
The agent will only drive an origin your workspace has verified. This is a hard boundary, not a setting: private and internal addresses are refused outright, and so is any origin you have not proven you control. Without it, a bug report would be a way to make our infrastructure issue requests at a third party.
Run allowances
Runs are metered. Each attempt counts, so a report checked three times uses three runs, and the bisect and counterfactual passes use additional runs on top. The allowance is shared: performed Takes draw from this same monthly agent-run pool, one take per run.
- Free — not available.
- Pro — 20 runs per month.
- Team — 50 runs per seat per month, pooled across the workspace.
- Enterprise — 500 runs per month.
What it is not good at
Being straight about this is cheaper than disappointing you: flows behind a login it cannot complete, anything needing a real payment, email or SMS step, bugs that depend on a specific device or on precise timing, drag-and-drop and canvas interactions, and anything where the page has changed enough since the recording that the original steps no longer exist. Most of these come back as blocked rather than as a wrong answer, which is the failure mode we chose.