Repository navigation
docs: test ROI framing for factory metrics (AR-17) - #890
warp-agent-staging[bot] wants to merge 2 commits into
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
This PR was generated with Warp. Comment |
|
Autoresearch: AR-20261009-17 is ready for review.
|
|
@warp-agent-staging[bot] I'm starting a first review of this pull request. You can view the conversation on Warp. I completed the review and no human review was requested for this pull request. Comment Powered by Oz |
There was a problem hiding this comment.
Review summary
The independent agent completed its review for this commit.
Findings
src/content/docs/factories/measure-and-improve.mdx:10— [SUGGESTION] The opening expands ROI as "return on investment" and says the metrics evaluate it, but the metrics are operational and Cost per PR is an estimate, so readers can take it as financial ROI. Requested change: state the scope in the sentence (for example "operational outcomes and estimated cost") or drop the expansion of ROI.src/content/docs/factories/measure-and-improve.mdx:21— [SUGGESTION] The paragraph under "Measuring agent ROI" restates the definitions of PRs opened, PRs merged, Autonomy, and PR cycle time that the metrics table defines a few lines below. Requested change: cut it to the baseline-comparison and run-count-is-not-shipped-work guidance and let the table carry the definitions.src/content/docs/factories/measure-and-improve.mdx:21— [SUGGESTION] The new section points readers at PRs merged, Autonomy, PR cycle time, and Cost per PR without the caveats further down: the first three need a connected code host, and Cost per PR is an estimate that can undercount. Requested change: add a short pointer to the caution and the Cost per PR note below, or state once in this paragraph that those metrics depend on a connected code host.
Verdict
Approve with nits
There was a problem hiding this comment.
Overview
This PR reframes the existing factory metrics overview around coding-agent ROI language without changing metric definitions or adding product behavior. I did not find security issues or spec drift; spec_context.md says no approved or repository spec context was found. The only concern is that the new ROI wording should stay clear that these are operational and cost signals, not a full financial ROI calculation.
Concerns
- The new description and opening copy use ROI phrasing in a way that could overpromise what the page measures; the inline suggestions narrow that framing while preserving the search term.
- Comments/tests audit: no code comments or tests were added or changed in this documentation-only diff.
Verdict
Found: 0 critical, 0 important, 2 suggestions
Approve with nits
Comment /warp-agent-review on this pull request to retrigger a review (up to 3 times on the same pull request).
Powered by Oz
| Measure coding agent ROI in Warp Factories with merged PRs, autonomy, | ||
| PR cycle time, cost per PR, and Scorers to evaluate completed runs. |
There was a problem hiding this comment.
💡 [SUGGESTION] The description promises measuring ROI, but the page documents operational and cost signals rather than a full ROI calculation. Keep the ROI keyword while narrowing the promise to signals.
| Measure coding agent ROI in Warp Factories with merged PRs, autonomy, | |
| PR cycle time, cost per PR, and Scorers to evaluate completed runs. | |
| Track coding agent ROI signals in Warp Factories with merged PRs, autonomy, | |
| PR cycle time, cost per PR, and Scorers to evaluate completed runs. |
| --- | ||
|
|
||
| Warp Factories tracks what your factory produces and how well it performs, so you can spot a problem, test a fix, and decide whether to keep it. | ||
| Warp Factories gives you metrics to evaluate the return on investment (ROI) of coding agents in a software factory: PRs opened and merged, autonomy, PR cycle time, and cost per PR. Use these alongside Scorers to assess completed runs, test a configuration change, and decide whether to keep it. |
There was a problem hiding this comment.
💡 [SUGGESTION] Expanding ROI as “return on investment” makes the sentence read like the dashboard calculates financial ROI. Narrow the claim to the signals the page actually documents.
| Warp Factories gives you metrics to evaluate the return on investment (ROI) of coding agents in a software factory: PRs opened and merged, autonomy, PR cycle time, and cost per PR. Use these alongside Scorers to assess completed runs, test a configuration change, and decide whether to keep it. | |
| Warp Factories gives you metrics to evaluate coding agent ROI signals in a software factory: PRs opened and merged, autonomy, PR cycle time, and cost per PR. Use these alongside Scorers to assess completed runs, compare results against a baseline, test a configuration change, and decide whether to keep it. |
|
@warp-agent-staging[bot] I'm starting a first review of this pull request. You can view the conversation on Warp. I completed the review and no human review was requested for this pull request. Comment Powered by Oz |
There was a problem hiding this comment.
Review summary
The independent agent completed its review for this commit.
Findings
src/content/docs/factories/measure-and-improve.mdx:21— [SUGGESTION] The new "Measuring agent ROI" paragraph restates the metric list from line 10 and the metric table, and its last sentence repeats the point at line 45 about run count versus shipped work. Requested change: delete the H2 and fold the one non-redundant idea (baseline comparison; opened vs. merged) into the intro or Dashboard section.src/content/docs/factories/measure-and-improve.mdx:21— [SUGGESTION] The paragraph presents PRs merged, Autonomy, and PR cycle time as available measures without saying they require a connected code host; the caution sits later at line 41. Requested change: add a short clause such as "(these need a connected code host; see the note below)" or move the caution before the metric guidance.src/content/docs/factories/measure-and-improve.mdx:10— [SUGGESTION] "Evaluate the return on investment (ROI) of coding agents" implies the Dashboard computes ROI, but it reports PR counts, autonomy, cycle time, and estimated cost per PR only. The terms are also inconsistent ("coding agent ROI" at line 10 and in the description, "agent ROI" in the H2 at line 19). Requested change: say the metrics measure outcomes and cost so readers can judge ROI themselves, and use one consistent term.src/content/docs/factories/measure-and-improve.mdx:4— [SUGGESTION] The new description covers only metrics and Scorers, but the page is titled "Measure and improve a factory" and also covers Benchmarks and Self-improvement. Requested change: keep ROI but name Benchmarks and Self-improvement, staying within 160 characters.src/content/docs/factories/measure-and-improve.mdx:21— [NIT] "Compare results against a baseline over the same date range" is ambiguous. Requested change: use "Compare results against a baseline period of equal length."
Verdict
Approve with nits
There was a problem hiding this comment.
Overview
This PR reframes the existing factory metrics overview around coding-agent ROI vocabulary by updating the meta description, opening paragraph, and first metrics section heading. I reviewed the docs copy, PR-body risk framing, changed-comment/test surface, and security impact; the diff preserves the existing metric definitions and adds no commands, executable code, UI paths, or security-sensitive instructions.
Concerns
- No blocking concerns found.
- No approved or repository spec context was provided for implementation drift validation.
- Security pass: no security findings; this is a public documentation copy change only.
Verdict
Found: 0 critical, 0 important, 0 suggestions
Approve
Comment /warp-agent-review on this pull request to retrigger a review (up to 3 times on the same pull request).
Powered by Oz
There was a problem hiding this comment.
Review summary
The independent agent completed its review for this commit.
Findings
src/content/docs/factories/measure-and-improve.mdx:10— [SUGGESTION] "what coding agent work ships" overstates what the Dashboard measures (PRs opened and merged, not deployments), and "not a calculation of financial return" is defensive phrasing. Requested change: say what the work "gets merged" and what it costs, and drop or soften the negation.src/content/docs/factories/measure-and-improve.mdx:19— [SUGGESTION] The new "## Measuring agent ROI" H2 holds one sentence and nests the unchanged metrics table under an H3. Requested change: fold the sentence into the existing "Read metrics on the Dashboard page" H2 at its original level, or give the new section content of its own.src/content/docs/factories/measure-and-improve.mdx:21— [SUGGESTION] "Compare these operational ROI signals against a baseline period of equal length" repeats line 10, doesn't say which metrics or how to choose the period, and duplicates the later "Collect a baseline" step; "ROI" is also repeated in the description, opening, heading, and this sentence. Requested change: delete the sentence or link to the "Collect a baseline" step, and reduce repeated "ROI" usage.
Verdict
Approve


Summary
Experiment: AR-20261009-17
Connect existing factory metrics to operational ROI vocabulary without claiming to calculate financial return or promising savings.
Hypothesis (unproven): naming ROI in the description, opening, and a descriptive H2 makes the documented metrics a closer semantic match for effectiveness queries and makes Warp attribution easier for answer engines.
Metric: Peec Agent Effectiveness & ROI topic, Sep 25–Oct 8, 2026; warp.dev retrieval 11/252 chats (4.4%), Warp visibility 0/252 (0%).
Prediction: retrieval ≥10% and visibility 3–8% within three weeks of merge.
Rollback: revert this PR's one-page copy change if the three-week result does not support the hypothesis.
Related issues
Sources supporting every product statement: existing metrics and caveats, Scorers, and benchmarks. No new product functionality, ROI calculator, or savings claim.
Content design plan
Worthiness: existing-page copy only; published metrics and configurable Scorers already have canonical documentation. No rollout or availability claim added.
Validation
How verified: npm typecheck (0 errors/warnings), heading-variable tests (5/5), build, heading TOCs (3,354 entries), homepage JSON-LD, and internal links (4,358 checked; 0 broken) passed. Committed changed-file style lint has warnings for existing metric names absent from glossary, no errors; UI-reference validation passes. Whitespace diff passes. Trunk is unavailable locally; CI remains the delivery gate. Pure copy change; no behavior regression test needed.
G2: PASS (no check score regresses). Public curl confirms the changed opening and ROI heading at https://docs-git-factory-ar-20261009-17-roi-warpdotdev.vercel.app/factories/measure-and-improve/.
npx -y @ora-ai/ax@0.5 audit https://docs-git-factory-ar-20261009-17-roi-warpdotdev.vercel.app --jsoncompleted; control 56/100 → PR 55/100. All 125 check scores are identical. The pricing-md and pricing-info na→fail flip is an applicability reclassification (also seen on unrelated docs #885), 0 points; aggregate denominator effect only. This preview-host classifier noise is documented, not attributed to copy or waived as a content regression. No production comparison; no feedback submission. Forced scans: 0.CI: Docs editorial quality, Docs technical references, Build/link-check/audit, CodeQL, and Vercel passed for this head. Editing this evidence can trigger another contract/build run.
Visual proof: Chromium preview verification shows legible wrapping with no clipped table cells; AR-20 pattern link also lands on its detail heading. Static copy screenshots are attached below; unrelated Vercel toolbar CSP warning is not a content error.
Control vs PR — all 125 checks
ard-catalogai-catalog-publishedard-entries-validard-trust-manifestagentic-search-usecaseagentic-search-specificbrand-search-accuracywikipedia-presencerobots-ai-policy-qualitymcp-registry-listednpm-sdk-packageagent-rules-repoagent-plugins-reporegistry-brandingchatgpt-app-listedsitemapcontent-no-jsbot-detectionagent-discovery-fileagent-skills-index-v2a2a-agent-cardskills-sh-listedpricing-mdnlweb-schema-feedsmcp-well-known-discoveryagent-mode-viewlink-headers-discoverymarkdown-url-fallbackmodular-llms-txtsitemap-lastmodrobots-agent-user-policyllms-txt-existsllms-txt-formattingjson-ldpricing-infopublic-api-docsagent-instructionskills-sh-qualityjson-ld-entity-linkingmetadata-completenessorg-schema-completenessschema-type-breadthtrust-anchorsllms-txt-links-resolvemarkdown-link-alternatemarkdown-frontmatterredirect-hygienepage-token-budgetcode-fence-validitydocs-auth-gateopenapi-specdeveloper-portalapi-catalog-rfc9727markdown-negotiationmarkdown-negotiation-varyagent-ua-markdownagent-crawler-reachabilitymcp-tool-descriptionsmcp-param-schemasmcp-server-identitymcp-tool-listingmcp-tool-namingpublic-apioauth-supportscoped-permissionsmcp-auth-mechanismmcp-oauth-metadatamcp-pkce-s256onboarding-frictionweb-bot-auth-directoryoauth-protected-resourceauth-md-existsauth-md-structureauth-md-walkthrough-simulationagent-auth-discovery-metadataagent-auth-www-authenticateagent-auth-endpoints-reachablemcp-servermcp-error-handlingmcp-transport-modernwebmcprate-limit-headersidempotency-key-supportjson-error-responsesapi-error-modelapi-versioning-policypagination-shapeasync-job-patterncli-toolrest-sdk-packagesnlweb-asknlweb-streamingresponse-schema-coveragemcp-tool-annotationsmcp-server-cardmcp-multi-surface-coveragesandbox-environmentbatch-endpointsmcp-resource-listinggraphql-error-type-definitiongraphql-versioning-policygraphql-pagination-patterngraphql-async-job-patterngraphql-schema-completenessgraphql-batch-mutationsagent-friendly-404ax-document-structureax-native-controlsax-accessible-namesax-form-labelingax-tree-injection-safemcp-app-registrya2ui-supportmcp-apps-ui-qualitymcp-view-domainmcp-view-cspapi-schema-analysisfunction-calling-compatmcp-resource-qualitympp-supportx402-supportucp-supportacp-supportacp-delegate-paymentap2-supportReview revision
Commit:
0f0cd97d8d5cca5440dff69826e6c209f72fbfbe. Narrowed the opening and description to operational ROI signals, explicitly not a calculation of financial return; trimmed repeated metric definitions and corrected equal-length baseline comparison. Description also covers Scorers, benchmarks, and Self-improvement.How verified: repeated local typecheck, heading-variable tests, build, rendered heading TOCs, homepage JSON-LD, internal links, changed-file style lint, and technical-reference validation passed. Public curl and browser inspection confirm this revised copy; screenshots labeled revision/revised supersede the original captures (unlabeled original AR-19/AR-20 captures retained as historical context only).
Re-ran the non-forced domain audit command: servedFromCache=True, scannedAt=2026-10-09T19:01:17.451+00:00; it still matches all 125 control scores, with the documented pricing applicability classification. This is cached domain-level G2 evidence, not a claim that ora freshly crawled the revised content. Public HTML/browser proof is fresh. No additional forced scan; total remains 0.
Follow-ups
OWNER-DECISION: Optional title “Measuring factory ROI: autonomy, PR cycle time, and cost per PR” (63 characters) proposed, not applied. The 50–160-character rule applies to descriptions; title convention also requires editorial judgment. Current title retained to keep the page's improvement purpose.
Documentation risk
Risk: low
Rationale: Product-meaning-preserving copy and search metadata; metric definitions, caveats, commands, UI labels, and billing behavior unchanged.
Source files consulted: src/content/docs/factories/measure-and-improve.mdx@main, src/content/docs/factories/measure-and-improve/scorers.mdx@main
Requested engineering reviewers: none
Engineering review status: not-applicable
Docs override: none
Unverified claims
None. Retrieval and visibility effects are hypotheses, not product claims.
Computer-use screenshots (10)
AR-19 triggers: Triggers overview opening and start of Available trigger types table
AR-19 triggers: remaining rows of Available trigger types table and Related pages
AR-20 title and opening paragraph on deployed Multi-agent orchestration page
Full Common patterns comparison table on deployed AR-20 page
AR-17 revision (opening): Measure and improve a factory page
AR-17 revision (ROI section): Measuring agent ROI and Dashboard metrics table, wrapping without clipping
AR-20 revision (opening): Multi-agent orchestration page top
AR-20 revision (Common patterns table): three-column table wraps legibly, no clipping
AR-19 revised opening and trigger table rows Slack through Schedules, including Jira Cloud and GitHub
AR-19 trigger table remaining rows (GitHub through Custom webhooks), completing all nine triggers