Layer Results
billing-only
Duration: 322s
Decision: GO
Go / No-Go
Deployment decision based on aggregated layer results. This page is the single decision document for the on-call engineer.
billing-only| Layer | Status | Key Metric | Reason |
|---|---|---|---|
| Service Health | Skip | — | Not included in this variant. |
| Pytest (CI) | Skip | — | Not included in this variant. |
| Data Quality | Skip | — | Not included in this variant. |
| AI QA | Skip | — | Not included in this variant. |
| SigNoz | Skip | — | Not included in this variant. |
| Billing Tests | Pass | 7/7 scenarios | All Stripe test-clock journeys green (282.9s). |
| Release Review | Pass | — | No failing CI, no open PRs. 0 repos with undeployed changes. |
Service Health
Endpoint health checks across staging and production environments.
Health checks were not included in this run variant.
Pytest (CI Tests)
GitHub Actions test workflows triggered across 0 repositories. Each repo's test.yml or ci.yml was dispatched and polled for completion.
Pytest CI triggers were not included in this run variant.
Forge
Forge-sentinel test suite and nightly runner workflows.
Forge-sentinel workflows were not included in this run variant.
Evals
Legacy pipeline evaluation suite.
Pipeline evals were not included in this run variant.
Forge Runner
MCP infrastructure and agent pipeline tests.
Parser V2
Document parsing smoke tests and E2E pipeline verification.
Content Evals
Parser output checked against ground truth on 15 documents (comparison-only).
30.7 vs 0.0, 13.5 vs 0.7893.<4.3 vs <3.4, 27.0-32.0 vs 0.0 - 0.0.pg vs g/dL, ∅ vs ratio.Serum vs Blood, Serum vs Blood.∅ vs Class 0-I, ∅ vs Class 0-I.91.7%. Lower precision = over-extraction.98.0%.94.7%.90.0%.| Predicted | |||
|---|---|---|---|
| No | Yes | ||
| Actual | No | 0true neg | 35false pos |
| Yes | 8false neg | 387true pos | |
| Field | Accuracy | Matches | Mismatches |
|---|---|---|---|
| result | 89.4% | 346 | 41 |
| reference range | 83.7% | 324 | 63 |
| unit | 82.4% | 319 | 68 |
| test date | 97.7% | 378 | 9 |
| sample source | 84.0% | 325 | 62 |
Report E2E
Layer 3.5 -- Report generation, validation, hallucination detection.
Playwright
End-to-end UX tests against the React frontend.
Playwright UX tests were not included in this run variant.
Agent Exploration
Automated route exploration to detect console errors, network failures, and performance issues.
Agent exploration was not included in this run variant.
Data Quality
Parser output validation — 17 checks across biomarkers, diagnoses, procedures, genetics. Results from staging Redis (last 48h). 0 records across 0 users.
Data validation was not included in this run variant.
SigNoz
Error and fatal log counts from SigNoz over the past 24 hours across all monitored services.
SigNoz observability checks were not included in this run variant.
Performance Baselines
Pre-deploy staging validation — checks whether anything has slowed down or started erroring more in staging before promoting develop to main. Production performance is monitored separately.
Performance baselines were not included in this run variant.
Stripe Test-Clock Billing Tests
variant: billing-onlySix time-shifted billing journeys driven by Stripe test clocks against staging: trial → first paid period, successful renewal, failed-renewal dunning, recovery from past_due, cancel-at-period-end, and plan-switch-at-next-cycle.
Assertion contract
billing-tests/CLAUDE.md/quota/status/billing on email_verified being true and no programmatic verify-skip exists. A CF-admin helper in authentication-service unblocks this.
active, starter pricestarterpast_dueactivecanceledfreeprofessionalprofessionalpracticepracticeRun locally
cd billing-tests && npm test. Requires billing-tests/.env.staging populated with BILLING_TESTS_STRIPE_TEST_KEY, BILLING_API_KEY, and the two plan price IDs. See billing-tests/CLAUDE.md.
AI QA Test Cases
Agent-driven test case execution via Claude Agent SDK + Playwright MCP.
Manual test cases were not included in this run variant.
CHR Reconciliation
Layer 5.9 -- Per-CHR lifecycle join across api-backend trigger logs, forge-sentinel terminal events, and the api-backend sweeper. Surfaces silent stalls, stuck reports, and clustered failures.
Release Review
Comparison of main vs develop across all tracked repositories.
Release review was not included in this run variant.