cargo ninety-nine
A cargo plugin for finding and tracking flaky tests in Rust projects using Bayesian inference.
Flaky tests – tests that sometimes pass and sometimes fail without code changes – erode trust in your test suite, waste CI time, and mask real regressions. cargo ninety-nine runs each test multiple times, applies Bayesian statistical analysis to compute a flakiness probability, and tracks results over time in a local SQLite database.
Key Features
- Bayesian flakiness detection – computes posterior probability of flakiness using Beta distribution, not just pass/fail ratios
- Filter DSL – expressive query language to target tests by name, package, binary, kind, or flakiness status
- Pattern analysis – detects time-of-day and environmental (CI vs local) failure patterns
- Trend tracking – monitors whether tests are improving, stable, or degrading over time
- Interactive TUI – terminal interface for browsing scores with sorting, filtering, and detail drill-down
- Quarantine management – manually or automatically quarantine flaky tests that exceed thresholds
- Multiple export formats – JUnit XML, HTML reports, CSV, and JSON for integration with other tools
- CI workflow generation – generates ready-to-use GitHub Actions and GitLab CI configurations
- Persistent storage – SQLite database with WAL mode for concurrent access, automatic data retention
Quick Example
# Initialize configuration
cargo ninety-nine init
# Run flaky test detection (10 iterations per test)
cargo ninety-nine test -n 10
# Run only flaky, non-quarantined tests
cargo ninety-nine test "flaky & !quarantined"
# Check a specific test's history
cargo ninety-nine status tests::my_flaky_test
# Export results as JSON
cargo ninety-nine export json results.json
How It Works
- Discover – compiles test binaries and lists all test cases
- Filter – applies filter expressions to select which tests to run
- Execute – runs each test N times with configurable concurrency and timeouts
- Analyze – applies Bayesian inference to compute P(flaky) with credible intervals
- Report – displays results with category labels (Stable, Occasional, Moderate, Frequent, Critical)
- Store – persists scores and run history for trend analysis across sessions
Getting Started
Prerequisites
- Rust 1.85+ (Edition 2024)
- cargo (included with Rust)
- Optionally, cargo-nextest for enhanced test running
Installation
Install from crates.io:
cargo install cargo-ninety-nine
Or build from source:
git clone https://github.com/glottologist/ninety-nine.git
cd ninety-nine
cargo install --path .
Verify Installation
cargo ninety-nine --version
Initialize a Project
Navigate to your Rust project and create a configuration file:
cd /path/to/your/project
cargo ninety-nine init
This creates a .ninety-nine.toml file with sensible defaults. You can overwrite an existing config with:
cargo ninety-nine init --force
First Test Run
Run detection with default settings (10 iterations per test):
cargo ninety-nine test
Filter to specific tests using a name pattern:
cargo ninety-nine test "my_module::tests"
Use the filter DSL to target specific subsets:
cargo ninety-nine test "flaky & !quarantined"
Increase iterations for higher confidence:
cargo ninety-nine test -n 50 --confidence 0.99
Output Formats
Add --output json to any command for machine-readable output:
cargo ninety-nine status --output json
cargo ninety-nine history --output json
Configuration
Configuration is stored in .ninety-nine.toml at the project root. All fields are optional — missing fields use defaults.
Generating a Config File
cargo ninety-nine init
Minimal Configuration
An empty .ninety-nine.toml file uses all defaults. You only need to specify values you want to change:
[detection]
min_runs = 20
confidence_threshold = 0.99
Full Configuration Reference
See Configuration Reference for all available options and their defaults.
Config Loading
The config file is loaded from the --project-dir path (defaults to .). If no config file exists, all default values are used. Unknown fields in the TOML file are silently ignored, so the config file is forward-compatible with newer versions.
Usage Overview
cargo ninety-nine provides seven subcommands:
| Command | Purpose |
|---|---|
test | Run tests repeatedly and compute flakiness scores |
init | Create a default configuration file |
status | View current flakiness scores and test detail |
history | View past detection sessions |
export | Export results to JUnit XML, HTML, CSV, or JSON |
quarantine | Manage test quarantine (list, add, remove) |
ci | Generate CI workflow files |
Interactive Mode
By default, test, status, and history launch an interactive TUI with sortable tables, category filtering, and detail drill-down. Pass --non-interactive (or -N) to use plain text output instead – useful for CI, scripts, or piped output.
Global Options
These options apply to all subcommands:
--project-dir <PATH> Project root directory (default: .)
--output <FORMAT> Output format: console or json (default: console)
--non-interactive, -N Disable TUI, use plain text output
--verbose, -v Verbose output during detection
Typical Workflow
- Initialize –
cargo ninety-nine init - Test –
cargo ninety-nine test -n 20 - Filter –
cargo ninety-nine test "flaky & !quarantined"(see Filter DSL) - Investigate –
cargo ninety-nine status tests::suspect_test - Quarantine –
cargo ninety-nine quarantine add tests::flaky_test --reason "timing-dependent" - Export –
cargo ninety-nine export html report.html - Automate –
cargo ninety-nine ci generate github
Running Tests
The test subcommand is the core functionality – it discovers, runs, and analyzes tests for flakiness.
Basic Usage
# Run all tests 10 times each (default)
cargo ninety-nine test
# Filter to specific tests by name pattern
cargo ninety-nine test "my_module::tests"
# 50 iterations with 99% confidence threshold
cargo ninety-nine test -n 50 --confidence 0.99
Options
| Option | Default | Description |
|---|---|---|
<filter_expr> | none | Filter expression – a test name regex or DSL expression |
-n, --iterations | 10 | Number of times to run each test |
--confidence | 0.95 | Confidence threshold for classifying a test as flaky |
Filter Expressions
The positional argument accepts either a plain regex pattern or a full filter DSL expression:
# Regex pattern matching test names
cargo ninety-nine test "my_module::tests"
# DSL: only tests marked flaky that are not quarantined
cargo ninety-nine test "flaky & !quarantined"
# DSL: tests in a specific package
cargo ninety-nine test "package(my_crate)"
# DSL: combine predicates
cargo ninety-nine test "test(network) & kind(test) & !quarantined"
See the Filter DSL guide for the full syntax reference.
What Happens During a Run
- Binary discovery – runs
cargo test --no-run --message-format json-render-diagnosticsto compile and locate test binaries - Test listing – executes each binary with
--list --format terseto enumerate test cases - Filtering – evaluates the filter expression (if provided) against each test’s metadata
- Parallel execution – runs each matched test individually N times using
--exact --nocaptureflags - Bayesian scoring – computes P(flaky) using Beta distribution with uniform prior
- Storage – writes results to SQLite database for trend tracking
- Reporting – displays results sorted by flakiness probability
Understanding the Output
Console Report
Flaky Test Detection Report
Test Runs Pass% P(flaky) Category
--------------------------------------------------------------------------------------------
tests::race_condition 10 60.0% 33.3% Frequent
tests::timing_dependent 10 80.0% 16.7% Moderate
tests::stable_test 10 100.0% 1.0% Stable
Flakiness Categories
| Category | P(flaky) Range | Meaning |
|---|---|---|
| Stable | < 1% | Reliably passes |
| Occasional | 1% - 5% | Rare failures, may be acceptable |
| Moderate | 5% - 15% | Notable flakiness, worth investigating |
| Frequent | 15% - 30% | Significant problem |
| Critical | > 30% | Severely broken, likely a bug |
Post-Run Analysis
After the main report, the tool displays:
- Detected Patterns – time-of-day or environmental correlations in failure data
- Degrading Trends – tests whose failure rate has increased between recent and previous runs
Auto-Quarantine
If quarantine.auto_quarantine = true in config, tests exceeding the configured thresholds are automatically quarantined after detection. See Quarantine Management.
Verbose Mode
Use -v for per-test progress output instead of the progress bar:
cargo ninety-nine test -v
Summary Mode
Set reporting.console.summary_only = true in config to show only aggregate counts:
Summary: 42 tests, 3 flaky, 39 stable
Multi-phase Diagnose
cargo ninety-nine diagnose finds flaky tests by stressing each test binary under multi-threaded load, then re-running only stress-failing candidates in serial isolation.
Phases
- Stress — for each binary that contains a selected test, run the full binary N times with
--test-threadsset (no--exact). Failures are parsed from libtest output and intersected with the selected set. - Isolation — each candidate is run alone (
--exact) N times, serially (concurrency 1, retries 0). - Classify — each candidate becomes one of:
- Contention — failed under stress, always passed alone
- Intrinsic — failed sometimes alone
- Broken — never passed alone
- Record (optional) — with
--record, Intrinsic failures may be re-run under rr when available (Linux soft dependency).
Contention meaning (V1)
V1 stress is intra-binary: each test binary is exercised multi-threaded as a whole. Cross-binary workspace races are out of scope until a later runner rewrite.
Usage
# Default config ([diagnose] in .ninety-nine.toml)
cargo ninety-nine diagnose
# Override run counts
cargo ninety-nine diagnose --stress-runs 5 --isolation-runs 20
# Filter + optional rr recording
cargo ninety-nine diagnose "pkg::" --record --record-dir .ninety-nine/recordings
# JSON for CI
cargo ninety-nine diagnose -N --output json
Configuration
[diagnose]
stress_runs = 3
isolation_runs = 10
stress_threads = 0 # 0 = host parallelism
stress_timeout_secs = 300
record = false
record_dir = ".ninety-nine/recordings"
record_attempts = 10
Identity
Diagnostic rows store package + binary + test name. Bayesian scores and test_runs still use the short test name so existing filters and the scores TUI keep working.
Soft CI exit
diagnose exits 0 after a successful run even when flaky classes are found. Fail the pipeline in a wrapper if you need a hard gate.
Interactive TUI
Without -N, diagnose opens a table of CLASS | STRESS | ISOLATION | TEST | REC.
| Key | Action |
|---|---|
j / k | Move |
f | Cycle class filter (all → contention → intrinsic → broken) |
| Enter | Detail overlay (counts + recording path) |
q | Quit |
rr chaos mode
cargo ninety-nine diagnose --record --chaos
Requires recording enabled (--record or diagnose.record = true). Passes --chaos to rr record.
Auto-quarantine by class
[quarantine]
enabled = true
auto_quarantine = true
[quarantine.by_class]
intrinsic = true
contention = false # default: leave load-sensitive tests visible
broken = true
Reasons are stored as auto:intrinsic, auto:broken, or auto:contention.
Multi-phase test
[detection]
multi_phase = false # default
cargo ninety-nine test --multi-phase
cargo ninety-nine test --no-multi-phase
When enabled: diagnose stress/isolation for candidates, then Bayesian multi-run only for non-candidates.
Platform matrix (rr)
| Platform | rr recording |
|---|---|
| Linux + rr on PATH | Available |
| Linux without rr | Soft skip + install hint |
| macOS / Windows | Soft skip (“not supported”) |
Interactive TUI
cargo ninety-nine includes a terminal user interface for browsing flakiness scores and session history. The TUI launches automatically when running status, history, or test in a terminal. Pass --non-interactive (or -N) to disable it.
Scores View
The scores view is the main screen after running test, or when running status without a test name argument. It uses bordered panels, a scrollbar, and colour-coded categories.
┌ Flaky Test Report ───────────────────────────────────────────┐
│ cargo ninety-nine | 42/42 tests shown │
└──────────────────────────────────────────────────────────────┘
Filter: All | Sort: P(flaky) (desc)
┌ Tests ──────────────────────────────────────────────────────┐^
│ Test Runs Pass% P(flaky) Cat. ││
│ ││
│ tests::network::retry_timeout 20 75.0% 0.250 Freq ││
│ tests::db::concurrent_writes 20 85.0% 0.150 Mod ││
│ tests::parser::edge_cases 20 95.0% 0.050 Occ │█
│ tests::math::addition 20 100.0% 0.010 Stab ││
│ ││
└──────────────────────────────────────────────────────────────┘v
j/k:nav s:sort r:reverse f:filter Enter:detail q:quit
Keybindings
| Key | Action |
|---|---|
j / Down arrow | Move selection down |
k / Up arrow | Move selection up |
s | Cycle sort field (Test, Runs, Pass%, P(flaky), Category) |
r | Reverse sort order |
f | Cycle category filter (All, Stable, Occasional, Moderate, Frequent, Critical) |
| Enter | Drill into selected test detail |
q / Esc | Quit |
| Ctrl+C | Quit |
Sorting
Press s to cycle through sort fields. Press r to toggle ascending/descending. The current sort field and direction are shown in the orange filter bar below the header.
Filtering
Press f to cycle through category filters. When a filter is active, only tests in that category are shown. The count in the header updates to reflect the filtered set.
Scrollbar
A vertical scrollbar appears on the right edge of the content panel. The scrollbar tracks the current selection position within the full list. The table viewport automatically follows the selected row, so scrolling through large lists works without manual page management.
Detail View
Press Enter on a test to open its detail overlay with a cyan border, showing:
- Score summary – category, P(flaky), confidence, pass/fail rates, total runs
- Bayesian parameters – alpha, beta, posterior mean, credible interval
- Trend – direction (Improving/Stable/Degrading) with score delta
- Failure patterns – correlated patterns with correlation percentage
- Recent runs – last 10 runs with outcome, duration, and timestamp
Press Enter, q, or Esc to return to the scores list.
History View
The history view shows past detection sessions with the same bordered-panel layout:
┌ Session History ─────────────────────────────────────────────┐
│ cargo ninety-nine | 13 sessions │
└──────────────────────────────────────────────────────────────┘
┌ Sessions ───────────────────────────────────────────────────┐^
│ Date Tests Flaky Branch Commit ││
│ ││
│ 2026-03-23 14:39 1803 0 jason/add_tests 92d65c74 │█
│ 2026-03-23 11:34 1803 0 jason/add_tests 92d65c74 ││
│ 2026-03-18 18:39 105 0 main 5ee51231 ││
└──────────────────────────────────────────────────────────────┘v
j/k:nav Enter:detail q:quit
Navigate with j/k or arrow keys. Press Enter to view the test runs from that session.
Session Detail
Pressing Enter on a session opens a scrollable overlay with the same filter, sort, and reverse controls available in the scores view:
┌────────── 2026-03-26 18:38 | jason/add_tests | 47dc1602 ──────────┐
│ 9580/9580 tests | 9580 passed | 0 failed │
│ Filter: All | Sort: Test (asc) │
│ Test Outcome Duration Retries ^│
│ ││
│ add_slots_test PASS 51ms 0 ││
│ address::tests::parse_case_1 PASS 50ms 0 █│
│ address::tests::parse_case_2 PASS 50ms 0 ││
│ address::tests::roundtrip PASS 50ms 0 ││
│ ││
│ j/k:nav s:sort r:reverse f:filter Enter/q/Esc:back v│
└────────────────────────────────────────────────────────────────────-┘
- Title bar – session date, branch, and commit hash
- Summary line – filtered/total tests, passed count, failed count
- Filter bar – current outcome filter and sort field with direction (same orange style as the scores view)
- Test table – test name, outcome (colour-coded PASS/FAIL/TIME/PANC/SKIP), duration, retry count
- Scrollbar – right edge of the test table with
^/vmarkers, tracks the selected row
Keybindings
| Key | Action |
|---|---|
j / Down arrow | Move selection down |
k / Up arrow | Move selection up |
s | Cycle sort field (Test, Outcome, Duration, Retries) |
r | Reverse sort order |
f | Cycle outcome filter (All, Pass, Fail, Timeout, Panic, Ignored) |
Enter / q / Esc | Return to session list |
| Ctrl+C | Quit |
Sorting
Press s to cycle through sort fields: Test name, Outcome, Duration, Retries. Press r to toggle ascending/descending. The current sort state is shown in the orange filter bar.
Filtering
Press f to cycle through outcome filters. When active, only runs with that outcome are shown. The summary line updates to show filtered/total counts.
Disabling the TUI
For CI pipelines, scripts, or piped output, use --non-interactive:
cargo ninety-nine status --non-interactive
cargo ninety-nine history --non-interactive
cargo ninety-nine test -n 10 --non-interactive
This produces the same text output as previous versions.
Visual Style
The TUI follows a panel-based layout inspired by tools like tOwl:
- Header panel – bordered with cyan, contains the tool name and summary statistics
- Filter bar – orange text showing current filter and sort state
- Content panel – bordered with dark grey, contains the data table with a title
- Scrollbar – right edge of content panel with
^/vend markers - Footer – keybinding hints with bold key names
- Category colours – Stable (green), Occasional (yellow), Moderate (red), Frequent (bold red), Critical (white on red)
Terminal Requirements
The TUI uses the alternate screen buffer and raw mode via crossterm. It works in any terminal emulator that supports ANSI escape sequences. On Unix, SIGTERM and SIGHUP trigger graceful shutdown. On all platforms, the terminal state is restored on exit, including after panics.
Minimum terminal size: 60 columns by 10 rows.
Filter DSL
The filter DSL lets you select which tests to run using an expressive query language. Filters are passed as the positional argument to the test command.
cargo ninety-nine test "<filter expression>"
Syntax Overview
A filter expression is built from predicates combined with boolean operators.
Bare Words
A bare word (any identifier that is not a keyword or function call) is treated as a regex pattern matched against the test name:
# Matches any test whose name contains "network"
cargo ninety-nine test "network"
# Regex: matches test names starting with "tests::db_"
cargo ninety-nine test "tests::db_"
Function Predicates
Function predicates filter tests by metadata fields. The argument is always a single identifier.
| Function | Matches |
|---|---|
test(pattern) | Test name matches the regex pattern |
package(name) | Test belongs to a package whose name contains name |
binary(name) | Test belongs to a binary whose name contains name |
kind(k) | Binary kind is k – one of lib, bin, test, example |
# Tests in the "my_crate" package
cargo ninety-nine test "package(my_crate)"
# Tests from test binaries only (excludes doctests, examples, etc.)
cargo ninety-nine test "kind(test)"
# Tests whose name matches a regex
cargo ninety-nine test "test(db_.*insert)"
Boolean Keywords
These keywords evaluate based on stored flakiness data:
| Keyword | Matches |
|---|---|
flaky | Tests with at least one recorded failure, P(flaky) > 1%, and confidence at or above the threshold |
quarantined | Tests currently in the quarantine list |
all | All tests (always true) |
# Only previously-detected flaky tests
cargo ninety-nine test "flaky"
# All quarantined tests
cargo ninety-nine test "quarantined"
Note: The
flakyandquarantinedkeywords rely on data from previous runs stored in the database. Runtestat least once before using them.
Operators
Combine predicates using boolean operators:
| Operator | Meaning | Example |
|---|---|---|
& | AND – both sides must match | flaky & kind(test) |
| | OR – either side must match | package(a) | package(b) |
! | NOT – inverts the predicate | !quarantined |
( ) | Grouping – controls evaluation order | (flaky | quarantined) & kind(test) |
Operator Precedence
! (NOT) binds tightest. & and | share equal precedence and associate left to right — unlike most programming languages, & does not bind tighter than |. Use parentheses whenever an expression mixes the two.
Example: flaky | quarantined & kind(test) is parsed as (flaky | quarantined) & kind(test), because the operators apply strictly left to right. To AND first, write flaky | (quarantined & kind(test)) explicitly.
Expressions may nest ! and parentheses at most 64 levels deep; anything deeper is rejected with a parse error.
Examples
Basic Filtering
# Run all tests (no filter)
cargo ninety-nine test
# Match test names by regex
cargo ninety-nine test "integration"
# Tests in a specific package
cargo ninety-nine test "package(my_lib)"
Combining Predicates
# Flaky tests that are not quarantined
cargo ninety-nine test "flaky & !quarantined"
# Tests from either of two packages
cargo ninety-nine test "package(auth) | package(session)"
# Network tests in the integration test binary
cargo ninety-nine test "test(network) & binary(integration)"
CI-Focused Patterns
# Re-run only known flaky tests with extra iterations
cargo ninety-nine test "flaky & !quarantined" -n 50
# Run test binaries only, skip examples and benches
cargo ninety-nine test "kind(test)"
# Quarantined tests only (verify if they are still flaky)
cargo ninety-nine test "quarantined" -n 30 --confidence 0.99
# Everything except quarantined tests
cargo ninety-nine test "!quarantined"
Complex Expressions
# Flaky tests in test binaries from a specific package
cargo ninety-nine test "flaky & kind(test) & package(core)"
# Either flaky or quarantined, but only from lib targets
cargo ninety-nine test "(flaky | quarantined) & kind(lib)"
Error Handling
Invalid filter expressions produce a clear error message:
$ cargo ninety-nine test "kind(invalid)"
# Error: unknown binary kind: invalid
$ cargo ninety-nine test "unknown_func(arg)"
# Error: unknown function: unknown_func
Valid kind values are: lib, bin, test, example.
Quarantine Management
Quarantine allows you to mark flaky tests for tracking and review. Quarantined tests are stored in the SQLite database.
List Quarantined Tests
cargo ninety-nine quarantine list
Output shows test name, flakiness score, whether it was auto-quarantined, and when:
Quarantined Tests
Test Score Auto Since
------------------------------------------------------------------------------------------------
tests::race_condition 33.3% yes 2026-03-05 14:30:00
tests::timing_test 20.0% no 2026-03-04 10:15:00
Add a Test to Quarantine
cargo ninety-nine quarantine add "tests::flaky_test" --reason "depends on network timing"
The --reason flag defaults to "manually quarantined" if omitted.
Remove a Test from Quarantine
cargo ninety-nine quarantine remove "tests::flaky_test"
Auto-Quarantine
Enable automatic quarantine in .ninety-nine.toml:
[quarantine]
enabled = true
auto_quarantine = true
[quarantine.threshold]
consecutive_failures = 3
failure_rate = 0.20
flakiness_score = 0.15
When auto_quarantine = true, any test detected as flaky (by the Bayesian detector) that also exceeds any of the three thresholds is automatically quarantined after a detect run.
Threshold Fields
| Field | Default | Description |
|---|---|---|
consecutive_failures | 3 | Number of consecutive failures at the tail of the run |
failure_rate | 0.20 | Overall failure rate (failures / total runs) |
flakiness_score | 0.15 | Bayesian P(flaky) posterior mean |
A test is auto-quarantined if is_flaky(score) AND (exceeds_score OR exceeds_failures OR exceeds_rate).
JSON Output
cargo ninety-nine quarantine list --output json
Exporting Results
Export flakiness scores to various file formats for integration with CI dashboards, issue trackers, or custom tooling.
Formats
JUnit XML
cargo ninety-nine export junit results.xml
Produces standard JUnit XML. Tests with P(flaky) >= 5% are marked as failures with details in the failure message. Compatible with CI systems that parse JUnit reports (GitHub Actions, GitLab, Jenkins, etc.).
HTML Report
cargo ninety-nine export html report.html
Generates a self-contained HTML page with a styled table showing all test scores. Categories are color-coded. Suitable for sharing with teams or archiving.
CSV
cargo ninety-nine export csv results.csv
Produces a CSV file with columns:
test_name,probability_flaky,pass_rate,total_runs,consecutive_failures,category,confidence
Values containing commas, quotes, or newlines are properly escaped per RFC 4180.
JSON
cargo ninety-nine export json results.json
Produces a JSON array of flakiness score objects with full detail, including Bayesian parameters. Useful for programmatic consumption, custom dashboards, or piping into other tools.
[
{
"test_name": "tests::example",
"probability_flaky": 0.167,
"confidence": 0.95,
"pass_rate": 0.8,
"fail_rate": 0.2,
"total_runs": 10,
"consecutive_failures": 1,
"bayesian_params": {
"alpha": 1.0,
"beta": 1.0,
"posterior_mean": 0.167,
"posterior_variance": 0.01,
"credible_interval_lower": 0.02,
"credible_interval_upper": 0.38
}
}
]
Data Source
Export uses the most recent flakiness scores stored in the SQLite database. Run test first to populate data.
CI Integration
Generating CI Workflows
Generate ready-to-use CI configuration files:
GitHub Actions
# Print to stdout
cargo ninety-nine ci generate github
# Write to file
cargo ninety-nine ci generate github .github/workflows/flaky-tests.yml
The generated workflow:
- Runs on a weekly schedule (Monday 3:00 AM UTC) and on manual dispatch
- Installs Rust, cargo-nextest, and cargo-ninety-nine
- Runs flaky test detection with your configured
min_runsandconfidence_threshold - Exports JUnit XML results
- Uploads results as a build artifact
GitLab CI
cargo ninety-nine ci generate gitlab .gitlab-ci-flaky.yml
The generated job:
- Runs on scheduled and manual pipelines
- Uses the
rust:latestDocker image - Installs cargo-nextest and cargo-ninety-nine
- Runs detection and exports JUnit XML
- Publishes JUnit artifacts for GitLab’s test report
Failure Behaviour
Generated workflows never fail the pipeline on flaky tests: the GitHub Actions job carries continue-on-error: true and the GitLab job carries allow_failure: true, so detection results arrive as reports rather than as red builds. Remove those lines from the generated file if you would rather have flaky detections break the build.
Manual CI Setup
If you prefer to configure CI manually:
# Install
cargo install cargo-ninety-nine
# Run detection
cargo ninety-nine test -n 20 --confidence 0.95
# Export for CI test report parsing
cargo ninety-nine export junit flaky-results.xml
Environment Detection
cargo ninety-nine automatically detects the CI environment:
| Environment Variable | Detected Provider |
|---|---|
GITHUB_ACTIONS | GitHub Actions |
GITLAB_CI | GitLab CI |
JENKINS_URL | Jenkins |
CIRCLECI | CircleCI |
BUILDKITE | Buildkite |
TF_BUILD | Azure DevOps |
The detected provider is stored with each test run, enabling environmental pattern analysis (CI vs local failure rate differences).
Best Practices
Guidance for getting the most out of cargo-ninety-nine in real-world workflows.
Iteration Strategy
Choosing min_runs
The number of iterations directly impacts detection accuracy:
| Iterations | Use Case | Confidence |
|---|---|---|
| 5–10 | Quick smoke check | Low — high false positive rate |
| 10–20 | Daily development | Moderate — good for known-flaky tests |
| 20–50 | Pre-merge gate | High — reliable classification |
| 50–100 | Baseline establishment | Very high — suitable for initial assessment |
Note: The Bayesian detector uses a uniform Beta(1,1) prior. With fewer than 10 runs, the prior dominates the posterior and scores are unreliable. The default
min_runs = 10is a minimum for meaningful results.
Building a Baseline
When first adopting cargo-ninety-nine, establish a baseline:
# Run with higher iterations to build confidence
cargo ninety-nine test -n 50
# Review the full status
cargo ninety-nine status
Subsequent runs build on the historical data, so daily runs with fewer iterations (10–20) are sufficient once the baseline exists.
Quarantine Strategy
When to Quarantine
Quarantine a test when:
- It blocks CI pipelines with intermittent failures
- It has a
Moderateor higher flakiness category (>= 0.05) - It has 3+ consecutive failures with no code changes
Do not quarantine a test when:
- It consistently fails — that is a real bug, not flakiness
- It just started failing after a recent change — investigate the change first
- The flakiness score is
Occasional(<0.05) — monitor it instead
Auto-Quarantine Thresholds
If enabling auto-quarantine, tune the thresholds to avoid over-quarantining:
[quarantine]
auto_quarantine = true
[quarantine.threshold]
consecutive_failures = 5 # stricter than default 3
failure_rate = 0.30 # stricter than default 0.20
flakiness_score = 0.20 # stricter than default 0.15
Start strict and relax gradually based on your team’s tolerance.
Quarantine Review Cadence
Quarantine entries persist until removed, so agree a review cadence with your team — fortnightly works well — and walk the list each time:
cargo ninety-nine quarantine list— see all quarantined tests- Re-run quarantined tests:
cargo ninety-nine test "quarantined()" - Remove fixed tests:
cargo ninety-nine quarantine remove <test_name>
CI Integration
Recommended CI Workflow
Run cargo-ninety-nine as a separate CI job that does not block merges:
# GitHub Actions example
flaky-detection:
runs-on: ubuntu-latest
if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- run: cargo install cargo-ninety-nine
- run: cargo ninety-nine test -n 20
- run: cargo ninety-nine export junit report.xml
- uses: actions/upload-artifact@v4
with:
name: flaky-report
path: report.xml
Warning: Avoid running flaky detection on every push. It adds significant CI time and produces noise. Schedule it nightly or weekly.
Shared Storage in CI
For teams wanting to accumulate results across CI runs, use PostgreSQL:
[storage]
backend = "Postgres"
[storage.postgres]
connection_string = "host=db.internal dbname=ninety_nine user=ci password=${NN_PG_PASSWORD}"
pool_size = 4
Pass the password via environment variable in CI secrets.
Performance Tuning
Parallel Execution
The parallel_runs setting controls how many tests execute concurrently:
[detection]
parallel_runs = 4 # default: 3
Guidelines:
- Local development: set to number of CPU cores / 2
- CI (shared runners): set to 1–2 to avoid resource contention
- Dedicated CI machines: set to CPU count
Reducing Test Discovery Time
Test discovery runs cargo test --no-run, which may trigger a full build. To speed this up:
- Keep your build cache warm (
target/directory) - Use incremental compilation
- Consider running detection on a subset:
cargo ninety-nine test "package(critical_module)"
Database Performance
SQLite (default):
- Excellent for single-machine use
- WAL mode is enabled automatically for concurrent reads
- Keep the database on a local SSD, not a network filesystem
PostgreSQL:
- Better for shared/team use and large datasets
- Tune
pool_sizebased on concurrent CI jobs - Use connection pooling (PgBouncer) for high-concurrency environments
Interpreting Results
Understanding Bayesian Scores
| Score | Meaning | Action |
|---|---|---|
| < 0.01 | Stable — no flakiness detected | No action needed |
| 0.01 – 0.05 | Occasional — rare failures | Monitor, investigate if trending up |
| 0.05 – 0.15 | Moderate — noticeable flakiness | Investigate root cause, consider quarantine |
| 0.15 – 0.30 | Frequent — regular failures | Fix or quarantine immediately |
| >= 0.30 | Critical — failing more than passing | Urgent fix required |
Acting on Patterns
| Pattern | Common Causes | Remediation |
|---|---|---|
| Time-of-day | Scheduled jobs competing for resources, time-sensitive assertions | Remove time dependencies, use faketime in tests |
| Environmental | Missing dependencies in CI, different OS behavior | Ensure CI mirrors local dev environment, use containers |
| Random | Race conditions, shared mutable state, non-deterministic ordering | Add synchronization, use deterministic seeds, isolate test state |
Confidence and Sample Size
Low confidence means the credible interval is wide — the true flakiness probability could be much higher or lower than the point estimate. Before acting on a score:
- High confidence (>0.95): Act on the score directly
- Moderate confidence (0.80–0.95): Run more iterations to confirm
- Low confidence (<0.80): Too few runs to draw conclusions, need more data
Migrating from SQLite to PostgreSQL
-
Export current data:
cargo ninety-nine export json current-data.json -
Update configuration:
[storage] backend = "Postgres" [storage.postgres] connection_string = "host=localhost dbname=ninety_nine" pool_size = 4 -
Run a fresh detection pass to populate the new database:
cargo ninety-nine test -n 20
Note: There is no automatic migration tool. Historical run data is not transferred — only flakiness scores and quarantine state can be preserved via JSON export. The new database will build fresh statistics from subsequent runs.
Team Workflows
Flaky Test Triage
Establish a regular triage process:
- Weekly: Review
cargo ninety-nine statusoutput - Per-sprint: Assign
Moderate+ flaky tests to developers - Per-release: Clear all
Frequent/Criticaltests before release
Filter Expressions for Triage
# Show only flaky tests
cargo ninety-nine test "flaky()"
# Focus on a specific package
cargo ninety-nine test "package(api) & flaky()"
# Exclude already-quarantined tests
cargo ninety-nine test "flaky() & !quarantined()"
Using JSON Output for Automation
# Export status as JSON for dashboards
cargo ninety-nine --output json status > flaky-status.json
# Parse with jq
cat flaky-status.json | jq '.[] | select(.probability_flaky > 0.1)'
Troubleshooting
Common issues and their solutions when using cargo-ninety-nine.
Installation Issues
no test runner available
error: no test runner available: install cargo-nextest or use cargo test
Cause: Neither cargo-nextest nor cargo test was found on the system PATH.
Solutions:
- Install cargo-nextest:
cargo install cargo-nextest - Verify Rust toolchain is installed:
rustup show - Ensure
~/.cargo/binis in your PATH
binary discovery failed
error: binary discovery failed: ...
Cause: cargo test --no-run failed to compile or list test binaries.
Solutions:
- Run
cargo test --no-runmanually to see the full compiler output - Fix any compilation errors in your project
- Ensure you are running from the project root (or use
--project-dir)
Configuration Issues
failed to parse config
error: failed to parse config: ...
Cause: The .ninety-nine.toml file contains invalid TOML syntax or unrecognized fields.
Solutions:
- Validate your TOML syntax: check for unclosed quotes, missing brackets, or incorrect indentation
- Re-generate a fresh config:
cargo ninety-nine init --force - Compare against the default config shown in the Configuration Reference
postgres backend selected but no config
error: invalid configuration: postgres backend selected but no [storage.postgres] config provided
Cause: storage.backend is set to "Postgres" but the [storage.postgres] section is missing.
Solution: Add the PostgreSQL configuration:
[storage]
backend = "Postgres"
[storage.postgres]
connection_string = "host=localhost dbname=ninety_nine user=postgres"
pool_size = 4
Test Execution Issues
Tests not discovered
Symptom: cargo ninety-nine test reports 0 tests found.
Causes and solutions:
- No test targets: Ensure your project has
#[test]functions or files intests/ - Filter too restrictive: Remove or widen your filter expression
- Wrong project directory: Use
--project-dir /path/to/project - Benchmark-only binaries: Only
#[test]functions are discovered, not benchmarks
Test timeouts
Symptom: Tests that normally pass are reported as Timeout.
Causes and solutions:
- Default timeout too low: The default is 300 seconds. For long-running tests, increase it in config:
[detection] # Timeout is controlled via the execution config - CI resource constraints: CI environments often have fewer resources. Check
memory_gbandcpu_countin the environment report - Test contention: Reduce
parallel_runsto lower resource contention:[detection] parallel_runs = 1
All tests show as flaky
Symptom: Every test gets a non-trivial flakiness score.
Causes and solutions:
- Too few iterations: With only a few runs, the Bayesian prior has outsized influence. Increase
min_runs:[detection] min_runs = 20 - Confidence threshold too low: Raise the threshold to require stronger evidence:
[detection] confidence_threshold = 0.99 - Systemic failures: If tests are failing due to environment issues (missing database, network), fix the root cause rather than tuning thresholds
Storage Issues
SQLite database locked
Symptom: storage error: database is locked
Causes and solutions:
- Concurrent access: Another
cargo-ninety-nineprocess may be running. WAL mode should handle most concurrent access, but heavy parallel writes can still lock - NFS/network filesystem: SQLite does not work reliably over network filesystems. Use a local path or switch to PostgreSQL
- Stale lock: If the process crashed, the lock file may remain. Delete
ninety-nine.db-walandninety-nine.db-shmnext to the database
PostgreSQL connection failures
Symptom: postgres storage error: connection refused or similar
Solutions:
- Verify PostgreSQL is running:
pg_isready - Check connection string format:
host=localhost port=5432 dbname=ninety_nine user=postgres password=... - Ensure the database exists:
createdb ninety_nine - Check network/firewall settings for remote connections
- Verify pool size is reasonable for your connection limits
Data retention and database size
Symptom: Database growing too large.
Solution: Configure retention_days to automatically purge old data:
[storage]
retention_days = 30 # default: 90
Data is purged at the end of each test run.
Filter DSL Issues
filter parse error
Symptom: error: filter parse error: ...
Common mistakes:
- Missing parentheses on predicates: Use
flaky()notflaky - Wrong operator syntax: Use
&for AND,|for OR,!for NOT - Unclosed parentheses: Ensure every
(has a matching) - Invalid regex in test(): The pattern must be a valid Rust regex
Valid examples:
flaky()
test(.*timeout.*)
package(auth) & !quarantined()
(flaky() | test(.*race.*)) & package(core)
CI Integration Issues
Workflow not triggering
Symptom: Generated CI workflow never runs.
Solutions:
- GitHub Actions: Ensure the workflow file is at
.github/workflows/. Check that scheduled triggers are on the default branch - GitLab CI: Ensure the pipeline is configured to run on schedules. Check that
rulesallow scheduled execution
CI environment not detected
Symptom: is_ci shows false in CI.
Cause: The CI provider’s environment variable is not set or not recognized.
Recognized variables:
| Variable | Provider |
|---|---|
GITHUB_ACTIONS | GitHub Actions |
GITLAB_CI | GitLab CI |
JENKINS_URL | Jenkins |
CIRCLECI | CircleCI |
TF_BUILD | Azure DevOps |
BUILDKITE | Buildkite |
Export Issues
Empty export files
Symptom: Export produces a file with no test data.
Cause: No flakiness scores have been computed yet.
Solution: Run cargo ninety-nine test at least once before exporting. Scores are computed and stored during the test run.
JUnit XML not recognized
Symptom: CI system does not parse the JUnit XML output.
Solution: Ensure the export path matches what your CI expects. For GitHub Actions:
- uses: dorny/test-reporter@v1
with:
artifact: ninety-nine-report
name: Flaky Tests
path: report.xml
reporter: java-junit
Getting More Information
Enable verbose output for detailed tracing:
cargo ninety-nine --verbose test
This sets the tracing subscriber to debug level, showing:
- Binary discovery details
- Test listing parsing
- Individual test execution results
- Storage operations
- Bayesian computation details
Types Reference
This chapter documents the core types used throughout cargo-ninety-nine.
Test Execution
TestRun
A single test execution record with full metadata.
#![allow(unused)]
fn main() {
pub struct TestRun {
pub id: Uuid,
pub test_name: TestName,
pub test_path: PathBuf,
pub outcome: TestOutcome,
pub duration: Duration,
pub timestamp: DateTime<Utc>,
pub commit_hash: String,
pub branch: String,
pub environment: TestEnvironment,
pub retry_count: u32,
pub error_message: Option<String>,
pub stack_trace: Option<String>,
}
}
| Field | Type | Description |
|---|---|---|
id | Uuid | Unique identifier for this run |
test_name | TestName | Type-safe test name |
test_path | PathBuf | Path to the test binary |
outcome | TestOutcome | Result of the execution |
duration | Duration | Wall-clock execution time |
timestamp | DateTime<Utc> | When the run occurred |
commit_hash | String | Git commit at time of run |
branch | String | Git branch at time of run |
environment | TestEnvironment | Execution environment details |
retry_count | u32 | Number of retries before this result |
error_message | Option<String> | Failure message, if any |
stack_trace | Option<String> | Stack trace on panic/failure |
TestOutcome
Classification of a test execution result.
#![allow(unused)]
fn main() {
pub enum TestOutcome {
Passed,
Failed,
Ignored,
Timeout,
Panic,
}
}
| Variant | Description |
|---|---|
Passed | Test completed successfully (exit code 0) |
Failed | Test assertion failed (non-zero exit, no panic) |
Ignored | Test marked with #[ignore] |
Timeout | Test exceeded the configured timeout |
Panic | Test panicked (detected via panicked at in output) |
Implements Display and FromStr for serialization. Display values: "passed", "failed", "ignored", "timeout", "panic".
TestEnvironment
Captures the environment where tests execute, used for pattern correlation.
#![allow(unused)]
fn main() {
pub struct TestEnvironment {
pub os: String,
pub rust_version: String,
pub cpu_count: u32,
pub memory_gb: f64,
pub is_ci: bool,
pub ci_provider: Option<String>,
}
}
Auto-detected at runtime from the host system. The is_ci field is inferred from environment variables (GITHUB_ACTIONS, GITLAB_CI, JENKINS_URL, CIRCLECI, TF_BUILD, BUILDKITE).
TestName
Newtype wrapper providing type-safe test names. Prevents accidental confusion with branch names, commit hashes, or other string fields.
#![allow(unused)]
fn main() {
pub struct TestName(String);
}
Conversions:
From<String>,From<&str>— construct from stringsDeref<Target = str>— borrow as&strAsRef<str>— reference conversionDisplay— format for outputPartialEq<str>,PartialEq<&str>— compare with string values
Methods:
into_inner(self) -> String— consume and return the inner string
Flakiness Detection
FlakinessScore
The primary output of the Bayesian detection engine. Contains the computed probability that a test is flaky along with statistical parameters.
#![allow(unused)]
fn main() {
pub struct FlakinessScore {
pub test_name: TestName,
pub probability_flaky: f64,
pub confidence: f64,
pub pass_rate: f64,
pub fail_rate: f64,
pub total_runs: u64,
pub consecutive_failures: u32,
pub last_updated: DateTime<Utc>,
pub bayesian_params: BayesianParams,
}
}
| Field | Type | Range | Description |
|---|---|---|---|
probability_flaky | f64 | [0.0, 1.0] | Posterior mean P(failure) |
confidence | f64 | [0.0, 1.0] | 1 - credible interval width (higher = more certain) |
pass_rate | f64 | [0.0, 1.0] | Fraction of runs that passed |
fail_rate | f64 | [0.0, 1.0] | Fraction of runs that failed (= 1 - pass_rate) |
total_runs | u64 | — | Number of non-ignored executions |
consecutive_failures | u32 | — | Trailing failure streak count |
BayesianParams
Full Bayesian computation state stored for auditability.
#![allow(unused)]
fn main() {
pub struct BayesianParams {
pub alpha: f64,
pub beta: f64,
pub posterior_mean: f64,
pub posterior_variance: f64,
pub credible_interval_lower: f64,
pub credible_interval_upper: f64,
}
}
| Field | Description |
|---|---|
alpha | Beta distribution shape parameter (prior + failures) |
beta | Beta distribution shape parameter (prior + passes) |
posterior_mean | alpha / (alpha + beta) |
posterior_variance | (alpha * beta) / ((alpha + beta)^2 * (alpha + beta + 1)) |
credible_interval_lower | 2.5th percentile of posterior Beta distribution |
credible_interval_upper | 97.5th percentile of posterior Beta distribution |
FlakinessCategory
Human-readable classification of flakiness severity.
#![allow(unused)]
fn main() {
pub enum FlakinessCategory {
Stable,
Occasional,
Moderate,
Frequent,
Critical,
}
}
| Category | Score Range | Console Color |
|---|---|---|
| Stable | < 0.01 | Green |
| Occasional | 0.01 – 0.05 | Yellow |
| Moderate | 0.05 – 0.15 | Orange |
| Frequent | 0.15 – 0.30 | Red |
| Critical | >= 0.30 | Dark Red |
Methods:
from_score(score: f64) -> Self— classify a probability valuelabel(&self) -> &'static str— human-readable label
Sessions
ActiveSession
Represents a running test session. Created via start(); to_run_session() produces the storable RunSession snapshot while the session stays live.
#![allow(unused)]
fn main() {
pub struct ActiveSession { /* private fields */ }
}
Methods:
| Method | Signature | Description |
|---|---|---|
start | fn start(commit_hash: &str, branch: &str) -> Self | Create a new session |
id | fn id(&self) -> &Uuid | Get session UUID |
to_run_session | fn to_run_session(&self) -> RunSession | Convert to storable form (borrowing) |
RunSession
A session record suitable for storage. May represent a running or completed session.
#![allow(unused)]
fn main() {
pub struct RunSession {
pub id: Uuid,
pub started_at: DateTime<Utc>,
pub finished_at: Option<DateTime<Utc>>,
pub test_count: u32,
pub flaky_count: u32,
pub commit_hash: String,
pub branch: String,
}
}
QuarantineEntry
A quarantined test record.
#![allow(unused)]
fn main() {
pub struct QuarantineEntry {
pub test_name: TestName,
pub quarantined_at: DateTime<Utc>,
pub reason: String,
pub flakiness_score: f64,
pub auto_quarantined: bool,
}
}
Analysis
TrendDirection
Direction of flakiness change over time.
#![allow(unused)]
fn main() {
pub enum TrendDirection {
Improving,
Stable,
Degrading,
}
}
A delta exceeding 0.05 (5%) triggers Improving (decreased flakiness) or Degrading (increased flakiness). Otherwise Stable.
TrendSummary
Trend analysis result comparing recent vs. historical flakiness.
#![allow(unused)]
fn main() {
pub struct TrendSummary {
pub test_name: TestName,
pub direction: TrendDirection,
pub recent_score: f64,
pub previous_score: f64,
pub score_delta: f64,
pub window_runs: u64,
}
}
FailurePattern
A detected failure pattern with correlation strength.
#![allow(unused)]
fn main() {
pub struct FailurePattern {
pub pattern_type: PatternType,
pub occurrences: u32,
pub correlation: f64,
pub examples: Vec<String>,
}
}
PatternType
Classification of detected failure patterns.
#![allow(unused)]
fn main() {
pub enum PatternType {
TimeOfDay,
Environmental,
Random,
}
}
| Variant | Trigger | Description |
|---|---|---|
TimeOfDay | Failure concentration > 3x expected in a specific hour | Failures cluster at a particular time |
Environmental | CI vs. local failure rate difference > 15% | Environment-specific failures |
Random | No pattern detected | Failures appear randomly distributed |
Bayesian Detector API
The BayesianDetector is the core statistical engine that computes flakiness probabilities using Beta-Binomial conjugate inference.
BayesianDetector
#![allow(unused)]
fn main() {
pub struct BayesianDetector {
prior_alpha: f64, // default: 1.0 (uniform prior)
prior_beta: f64, // default: 1.0 (uniform prior)
confidence_threshold: f64,
}
}
Constructor
#![allow(unused)]
fn main() {
pub const fn new(confidence_threshold: f64) -> Self
}
Creates a detector with a uniform Beta(1, 1) prior and the given confidence threshold.
| Parameter | Type | Description |
|---|---|---|
confidence_threshold | f64 | Minimum confidence required to classify a test as flaky (typically 0.95) |
Methods
calculate_flakiness_score
#![allow(unused)]
fn main() {
pub fn calculate_flakiness_score(
&self,
test_name: &str,
runs: &[TestRun],
) -> FlakinessScore
}
Computes a FlakinessScore from a set of test runs using Bayesian inference.
Algorithm:
- Count passes and failures from
runs(ignoringIgnoredoutcomes;Failed,Panic, andTimeoutall count as failures) - Update the Beta prior:
alpha = prior_alpha + failures,beta = prior_beta + passes - Compute posterior mean:
alpha / (alpha + beta) - Compute posterior variance:
(alpha * beta) / ((alpha + beta)^2 * (alpha + beta + 1)) - Compute 95% credible interval via
Beta::inverse_cdf(0.025)andBeta::inverse_cdf(0.975) - Confidence =
1.0 - (upper - lower)(narrower interval = higher confidence) - Count consecutive trailing failures
Returns: A FlakinessScore with all computed fields populated.
is_flaky
#![allow(unused)]
fn main() {
pub fn is_flaky(&self, score: &FlakinessScore) -> bool
}
Determines whether a test should be classified as flaky.
Criteria: Returns true when both conditions are met:
score.probability_flaky > 0.01— non-trivial failure probabilityscore.confidence >= confidence_threshold— sufficient statistical certainty
Internal Functions
These are not public but documented for understanding the algorithm:
| Function | Purpose |
|---|---|
count_outcomes(runs) | Counts passes and failures, ignoring Ignored outcomes |
count_consecutive_trailing_failures(runs) | Counts failures from the end of the run list until a pass is encountered |
credible_interval(alpha, beta) | Computes 2.5th–97.5th percentile interval using the statrs Beta distribution |
Usage Example
#![allow(unused)]
fn main() {
use cargo_ninety_nine::detector::BayesianDetector;
let detector = BayesianDetector::new(0.95);
let score = detector.calculate_flakiness_score("tests::my_test", &runs);
if detector.is_flaky(&score) {
println!("{} is flaky (P={:.2})", score.test_name, score.probability_flaky);
}
}
Statistical Properties
The detector guarantees the following properties (verified by property-based tests):
probability_flakyis always in[0.0, 1.0]confidenceis always in[0.0, 1.0]- More failures relative to passes always produces a higher
probability_flaky - The credible interval is always valid:
lower >= 0.0,upper <= 1.0,lower <= upper
See the Bayesian Detection reference for the full mathematical model.
Analysis API
The analysis module provides three capabilities: failure rate computation, trend detection, and failure pattern recognition.
Failure Rate
failure_rate
#![allow(unused)]
fn main() {
pub fn failure_rate(runs: &[&TestRun]) -> f64
}
Computes the fraction of runs that resulted in a failure outcome.
| Outcome | Counted as |
|---|---|
Passed | Non-failure |
Ignored | Non-failure |
Failed | Failure |
Panic | Failure |
Timeout | Failure |
Returns: A value in [0.0, 1.0]. Returns 0.0 for empty input.
Trend Detection
calculate_trend
#![allow(unused)]
fn main() {
pub fn calculate_trend(
test_name: &str,
runs: &[TestRun],
window: u32,
) -> Option<TrendSummary>
}
Analyzes the direction of flakiness change by comparing recent runs against historical runs.
| Parameter | Type | Description |
|---|---|---|
test_name | &str | Name of the test being analyzed |
runs | &[TestRun] | Runs to analyze (newest first not required) |
window | u32 | Maximum number of runs to consider |
Algorithm:
- Takes up to
windowmost recent runs - Requires a minimum of 4 runs; returns
Noneotherwise - Splits runs at the midpoint into “recent” (first half) and “previous” (second half)
- Computes failure rate for each half
- Classifies direction based on the delta:
- Delta < -0.05 (5% improvement):
Improving - Delta > +0.05 (5% regression):
Degrading - Otherwise:
Stable
- Delta < -0.05 (5% improvement):
Returns: Some(TrendSummary) with direction, scores, and delta, or None if insufficient data.
Duration Regression
detect_duration_regressions
#![allow(unused)]
fn main() {
pub fn detect_duration_regressions(
test_name: &str,
runs: &[TestRun],
min_history: usize,
threshold: RegressionThreshold,
) -> Option<DurationRegression>
}
Detects if the most recent test run is significantly slower than historical runs.
| Parameter | Type | Description |
|---|---|---|
test_name | &str | Name of the test |
runs | &[TestRun] | Runs ordered newest first (at least min_history required) |
min_history | usize | Minimum number of runs required |
threshold | RegressionThreshold | How the latest duration is judged against the baseline |
#![allow(unused)]
fn main() {
pub enum RegressionThreshold {
StdDevs(f64),
Multiplier(f64),
}
}
Algorithm:
- Returns
Noneif fewer thanmin_historyruns - Computes mean and standard deviation of the historical durations (the latest run is excluded from the baseline)
- Returns
Noneif the effective standard deviation is near zero StdDevs(z)flags when the latest duration exceedsmean + z × std_dev;Multiplier(m)flags when it exceedsmean × m- Returns
DurationRegressionwith deviation statistics
DurationRegression
#![allow(unused)]
fn main() {
pub struct DurationRegression {
pub test_name: TestName,
pub current_ms: f64,
pub mean_ms: f64,
pub std_dev_ms: f64,
pub deviation_factor: f64,
}
}
| Field | Description |
|---|---|
current_ms | Duration of the latest run in milliseconds |
mean_ms | Historical mean duration in milliseconds |
std_dev_ms | Historical standard deviation in milliseconds |
deviation_factor | How many standard deviations above mean: (current - mean) / std_dev |
Pattern Detection
detect_patterns
#![allow(unused)]
fn main() {
pub fn detect_patterns(runs: &[TestRun]) -> Vec<FailurePattern>
}
Scans test runs for recurring failure patterns. Returns all detected patterns.
Detection strategies:
Time-of-Day Pattern
Bins all failures by hour (0–23). If any hour contains more than 3x the expected random concentration, a TimeOfDay pattern is reported.
- Correlation value: ratio of peak-hour failures to total failures
- Examples: formatted as
"Hour HH: N failures"for the peak hour
Environmental Pattern
Compares failure rates between CI and local environments. If the difference exceeds 15%, an Environmental pattern is reported.
- Correlation value: absolute difference between CI and local failure rates
- Examples: formatted as
"CI failure rate: X%, local: Y%"
Random Fallback
If no time-of-day or environmental pattern is detected, a Random pattern is returned with correlation 0.0, indicating failures appear uniformly distributed.
Usage Example
#![allow(unused)]
fn main() {
use cargo_ninety_nine::analysis::{calculate_trend, detect_patterns};
use cargo_ninety_nine::analysis::duration::detect_duration_regressions;
// Trend analysis
if let Some(trend) = calculate_trend("tests::my_test", &runs, 100) {
println!("Trend: {} (delta: {:.2})", trend.direction, trend.score_delta);
}
// Duration regression
use cargo_ninety_nine::analysis::duration::RegressionThreshold;
if let Some(reg) =
detect_duration_regressions("tests::my_test", &runs, 5, RegressionThreshold::StdDevs(2.0))
{
println!("SLOW: {}ms vs mean {}ms ({:.1}x std_dev)",
reg.current_ms, reg.mean_ms, reg.deviation_factor);
}
// Pattern detection
for pattern in detect_patterns(&runs) {
println!("Pattern: {} (correlation: {:.2})", pattern.pattern_type, pattern.correlation);
}
}
Filter DSL API
The filter module implements a small domain-specific language for selecting tests. It consists of four stages: tokenization, parsing, context building, and evaluation.
Pipeline Overview
input string → tokenize() → parse() → FilterExpr (AST)
↓
test metadata + EvalContext → eval() → bool (match/no-match)
Compilation
compile_filter
#![allow(unused)]
fn main() {
pub fn compile_filter(input: &str) -> Result<FilterExpr, NinetyNineError>
}
Compiles a filter expression string into an AST. This is the primary entry point for the filter DSL.
Errors: Returns NinetyNineError::FilterParse if the input is syntactically invalid.
build_eval_context
#![allow(unused)]
fn main() {
pub async fn build_eval_context(
storage: &impl Storage,
confidence: f64,
) -> Result<EvalContext, NinetyNineError>
}
Pre-loads data from storage needed to evaluate predicates like flaky() and quarantined().
The context contains:
- Flaky test set: tests where
probability_flaky > 0.01andconfidence >= threshold - Quarantined test set: all currently quarantined tests
AST Types
FilterExpr
The abstract syntax tree for filter expressions.
#![allow(unused)]
fn main() {
pub enum FilterExpr {
And(Vec<FilterExpr>),
Or(Vec<FilterExpr>),
Not(Box<FilterExpr>),
Predicate(Predicate),
}
}
Predicate
Leaf-level predicates that match against test metadata.
#![allow(unused)]
fn main() {
pub enum Predicate {
Test(Regex), // match test name against regex
Package(String), // match package name
Binary(String), // match binary name
Kind(BinaryKind), // match binary kind (lib, bin, test, example)
Flaky, // test is currently flaky
Quarantined, // test is currently quarantined
All, // matches everything
}
}
Tokenizer
tokenize
#![allow(unused)]
fn main() {
pub fn tokenize(input: &str) -> Result<Vec<Token>, NinetyNineError>
}
Splits input into tokens.
Token
#![allow(unused)]
fn main() {
pub enum Token {
Ident(String),
LParen,
RParen,
And,
Or,
Not,
Equals,
}
}
Parser
parse
#![allow(unused)]
fn main() {
pub fn parse(tokens: Vec<Token>) -> Result<FilterExpr, NinetyNineError>
}
Recursive descent parser with this precedence (lowest to highest):
- Or (
|) — binary, left-associative - And (
&) — binary, left-associative - Not (
!) — unary prefix - Primary — parenthesized expressions or predicate calls
Predicate syntax:
test(pattern) — regex match on test name
package(name) — exact match on package name
binary(name) — exact match on binary name
kind(lib|bin|test|example) — match binary kind
flaky() — test is flaky
quarantined() — test is quarantined
all() — match everything
Evaluator
eval
#![allow(unused)]
fn main() {
pub fn eval(
expr: &FilterExpr,
meta: &TestMetadata,
ctx: &EvalContext,
) -> bool
}
Evaluates a compiled FilterExpr against a test’s metadata using a pre-built evaluation context.
TestMetadata
#![allow(unused)]
fn main() {
pub struct TestMetadata {
pub name: String,
pub package_name: String,
pub binary_name: String,
pub kind: BinaryKind,
}
}
Constructed from a TestCase at evaluation time.
EvalContext
#![allow(unused)]
fn main() {
pub struct EvalContext {
pub flaky_tests: HashSet<String>,
pub quarantined_tests: HashSet<String>,
}
}
Pre-computed sets for O(1) predicate evaluation.
Examples
# All flaky tests in the "auth" package
package(auth) & flaky()
# Everything except quarantined tests
!quarantined()
# Tests matching a pattern OR known flaky tests
test(.*timeout.*) | flaky()
# Complex composition
(package(api) | package(core)) & !quarantined() & flaky()
See the Filter DSL Guide for usage examples from the CLI.
Storage API
The storage module provides an async trait abstraction over test result persistence, with SQLite and PostgreSQL backends.
Storage Trait
#![allow(unused)]
fn main() {
pub trait Storage: Send + Sync {
async fn store_session(&self, session: &RunSession) -> Result<(), NinetyNineError>;
async fn finish_session(&self, session_id: &Uuid, test_count: u32, flaky_count: u32) -> Result<(), NinetyNineError>;
async fn store_test_run(&self, run: &TestRun, session_id: &Uuid) -> Result<(), NinetyNineError>;
async fn store_flakiness_score(&self, score: &FlakinessScore) -> Result<(), NinetyNineError>;
async fn get_test_runs(&self, test_name: &str, limit: u32) -> Result<Vec<TestRun>, NinetyNineError>;
async fn get_recent_sessions(&self, limit: u32) -> Result<Vec<RunSession>, NinetyNineError>;
async fn get_all_scores(&self) -> Result<Vec<FlakinessScore>, NinetyNineError>;
async fn get_score(&self, test_name: &str) -> Result<Option<FlakinessScore>, NinetyNineError>;
async fn quarantine_test(&self, test_name: &str, reason: &str, score: f64, auto: bool) -> Result<(), NinetyNineError>;
async fn unquarantine_test(&self, test_name: &str) -> Result<(), NinetyNineError>;
async fn get_quarantined_tests(&self) -> Result<Vec<QuarantineEntry>, NinetyNineError>;
async fn is_quarantined(&self, test_name: &str) -> Result<bool, NinetyNineError>;
async fn get_session_runs(&self, session_id: &Uuid) -> Result<Vec<TestRun>, NinetyNineError>;
async fn purge_older_than(&self, days: u32) -> Result<u64, NinetyNineError>;
}
}
All methods are async to support both synchronous (SQLite via spawn_blocking) and natively async (PostgreSQL) backends.
Method Reference
| Method | Description |
|---|---|
store_session | Persists a new run session |
finish_session | Marks a session as complete with final test/flaky counts |
store_test_run | Stores a single test execution result |
store_flakiness_score | Upserts a computed flakiness score (INSERT OR REPLACE) |
get_test_runs | Retrieves runs for a test, most recent first |
get_recent_sessions | Lists recent sessions, most recent first |
get_all_scores | Returns all scores ordered by probability_flaky descending |
get_score | Returns the score for a specific test, or None |
quarantine_test | Adds a test to quarantine with reason and score |
unquarantine_test | Removes a test from quarantine |
get_quarantined_tests | Lists all quarantined tests |
is_quarantined | Checks if a specific test is quarantined |
get_session_runs | Retrieves all test runs for a given session, ordered by test name |
purge_older_than | Deletes test runs older than N days, returns count deleted |
StorageBackend Enum
#![allow(unused)]
fn main() {
pub enum StorageBackend {
Sqlite(SqliteStorage),
Postgres(PostgresStorage),
}
}
Delegates all Storage trait methods to the underlying backend.
Factory Function
open_storage
#![allow(unused)]
fn main() {
pub async fn open_storage(config: &Config) -> Result<StorageBackend, NinetyNineError>
}
Opens the configured storage backend:
StorageBackendType::Sqlite— Opens (or creates) a SQLite database at the configured path, defaulting to$XDG_DATA_HOME/ninety-nine/ninety-nine.dbStorageBackendType::Postgres— Connects to PostgreSQL using the configured connection string and pool size
Errors: Returns InvalidConfig if Postgres is selected but no [storage.postgres] configuration is provided.
SQLite Backend
SqliteStorage
#![allow(unused)]
fn main() {
pub struct SqliteStorage { /* private fields */ }
}
Constructor:
#![allow(unused)]
fn main() {
pub fn open(db_path: &Path) -> Result<Self, NinetyNineError>
}
Opens the database, creates parent directories if needed, enables WAL mode for concurrent reads/writes, and runs schema migrations.
Features:
- WAL (Write-Ahead Logging) mode for better concurrent access
- Automatic schema creation on first open
- Bundled SQLite via
rusqlite(no system dependency required)
Default location: $XDG_DATA_HOME/ninety-nine/ninety-nine.db
- Linux:
~/.local/share/ninety-nine/ninety-nine.db - macOS:
~/Library/Application Support/ninety-nine/ninety-nine.db
PostgreSQL Backend
PostgresStorage
#![allow(unused)]
fn main() {
pub struct PostgresStorage { /* private fields */ }
}
Constructor:
#![allow(unused)]
fn main() {
pub async fn connect(
connection_string: &str,
pool_size: u32,
) -> Result<Self, NinetyNineError>
}
Connects to PostgreSQL and initializes the schema. Uses deadpool-postgres for connection pooling.
| Parameter | Type | Description |
|---|---|---|
connection_string | &str | PostgreSQL connection URL |
pool_size | u32 | Maximum number of pooled connections |
Features:
- Connection pooling via
deadpool-postgres - Automatic schema creation on first connect
- Natively async operations (no
spawn_blockingneeded)
Database Schema
Both backends share the same logical schema with four tables:
-- Test execution sessions
CREATE TABLE sessions (
id TEXT PRIMARY KEY,
started_at TEXT NOT NULL,
finished_at TEXT,
test_count INTEGER NOT NULL DEFAULT 0,
flaky_count INTEGER NOT NULL DEFAULT 0,
commit_hash TEXT NOT NULL,
branch TEXT NOT NULL
);
-- Individual test run results
CREATE TABLE test_runs (
id TEXT PRIMARY KEY,
session_id TEXT NOT NULL REFERENCES sessions(id),
test_name TEXT NOT NULL,
test_path TEXT NOT NULL,
outcome TEXT NOT NULL,
duration_ms INTEGER NOT NULL,
timestamp TEXT NOT NULL,
commit_hash TEXT NOT NULL,
branch TEXT NOT NULL,
environment TEXT NOT NULL,
retry_count INTEGER NOT NULL DEFAULT 0,
error_message TEXT,
stack_trace TEXT
);
-- Computed flakiness scores (upserted)
CREATE TABLE flakiness_scores (
test_name TEXT PRIMARY KEY,
probability_flaky REAL NOT NULL,
confidence REAL NOT NULL,
pass_rate REAL NOT NULL,
fail_rate REAL NOT NULL,
total_runs INTEGER NOT NULL,
consecutive_failures INTEGER NOT NULL DEFAULT 0,
last_updated TEXT NOT NULL,
bayesian_params TEXT NOT NULL
);
-- Quarantine entries
CREATE TABLE quarantine (
test_name TEXT PRIMARY KEY,
quarantined_at TEXT NOT NULL,
reason TEXT NOT NULL,
flakiness_score REAL NOT NULL DEFAULT 0.0,
auto_quarantined INTEGER NOT NULL DEFAULT 0
);
Utility Functions
| Function | Signature | Description |
|---|---|---|
parse_timestamp | fn(s: &str) -> DateTime<Utc> | Parses RFC 3339 timestamps, falls back to Utc::now() |
duration_to_ms | fn(d: Duration) -> i64 | Converts Duration to milliseconds for storage |
ms_to_duration | fn(ms: i64) -> Duration | Converts milliseconds back to Duration |
See the Storage Reference for configuration and migration details.
Runner API
The runner module handles test discovery, execution, and result collection.
Test Discovery
discover_test_binaries
#![allow(unused)]
fn main() {
pub fn discover_test_binaries(
project_root: &Path,
) -> Result<Vec<TestBinary>, NinetyNineError>
}
Discovers all test binaries in a Cargo project by running cargo test --no-run --message-format json-render-diagnostics and parsing the output.
Returns: A list of TestBinary structs, each containing the binary path, package name, and kind.
Errors: Returns BinaryDiscovery if cargo fails or output cannot be parsed.
list_tests_parallel
#![allow(unused)]
fn main() {
pub async fn list_tests_parallel(
binaries: &[TestBinary],
concurrency: usize,
) -> Result<Vec<TestCase>, NinetyNineError>
}
Lists all tests across multiple binaries concurrently. Each binary is invoked with --list --format terse and output lines ending with : test or : benchmark are parsed.
Uses tokio::sync::Semaphore for concurrency control.
Errors: Returns TestListing if a binary fails to produce test listings.
cargo_available
#![allow(unused)]
fn main() {
pub fn cargo_available() -> bool
}
Returns whether cargo is on PATH. The native runner builds test binaries via cargo test --no-run and executes them directly, so cargo is the only external tool required.
Test Execution
Executor
#![allow(unused)]
fn main() {
pub struct Executor<'a> {
config: &'a ExecutionConfig,
}
}
Runs individual test cases with retry support.
Constructor:
#![allow(unused)]
fn main() {
pub fn new(config: &'a ExecutionConfig) -> Self
}
run_single
#![allow(unused)]
fn main() {
pub fn run_single(
&self,
test_case: &TestCase,
) -> Result<TestResult, NinetyNineError>
}
Executes a single test case with retries. Spawns the test binary with the --exact flag targeting the specific test.
Retry behavior:
- Retries up to
config.retriestimes on failure - Stops immediately on first pass
- Applies
config.retry_delaybetween attempts
Timeout: Uses polling-based detection (50ms intervals). Kills the process when the deadline is exceeded, returning TestOutcome::Timeout.
Outcome classification:
| Condition | Outcome |
|---|---|
| Exit code 0 | Passed |
panicked at in stderr/stdout | Panic |
| Deadline exceeded | Timeout |
| Other non-zero exit | Failed |
ExecutionConfig
#![allow(unused)]
fn main() {
pub struct ExecutionConfig {
pub concurrency: usize,
pub timeout: Duration,
pub retries: u32,
pub retry_delay: Duration,
}
}
| Field | Default | Description |
|---|---|---|
concurrency | — | Maximum parallel test binary invocations |
timeout | 300s | Per-test execution timeout |
retries | 0 | Number of retry attempts on failure |
retry_delay | 100ms | Delay between retry attempts |
TestResult
#![allow(unused)]
fn main() {
pub struct TestResult {
pub test_case: TestCase,
pub outcome: TestOutcome,
pub duration: Duration,
pub stdout: String,
pub stderr: String,
pub attempt: u32,
}
}
Test Case Types
TestCase
#![allow(unused)]
fn main() {
pub struct TestCase {
pub name: TestName,
pub binary_path: PathBuf,
pub binary_name: String,
pub package_name: String,
pub binary_kind: BinaryKind,
pub kind: TestKind,
}
}
TestKind
#![allow(unused)]
fn main() {
pub enum TestKind {
Test,
Benchmark,
}
}
BinaryKind
#![allow(unused)]
fn main() {
pub enum BinaryKind {
Lib,
Bin,
Test,
Example,
}
}
Derived from Cargo metadata target kinds.
TestBinary
#![allow(unused)]
fn main() {
pub struct TestBinary {
pub path: PathBuf,
pub package_name: String,
pub binary_name: String,
pub kind: BinaryKind,
}
}
High-Level Runner
NativeRunner
#![allow(unused)]
fn main() {
pub struct NativeRunner { /* private fields */ }
}
Constructor:
#![allow(unused)]
fn main() {
pub fn new(project_root: &Path, config: ExecutionConfig) -> Self
}
Methods:
| Method | Description |
|---|---|
discover_tests(&self, filter: &str) | Discovers test cases, optionally filtered by name substring |
RunnerBackend
#![allow(unused)]
fn main() {
pub enum RunnerBackend {
Native(NativeRunner),
}
}
Extensible enum wrapping runner implementations. Currently supports native Cargo test execution.
Methods: native(), execution_config(), discover_tests() — all delegate to the inner NativeRunner.
Standalone Function
execute_iterations
#![allow(unused)]
fn main() {
pub fn execute_iterations(
test_case: &TestCase,
iterations: u32,
config: &ExecutionConfig,
environment: &TestEnvironment,
) -> Result<Vec<TestRun>, NinetyNineError>
}
Convenience function that runs a test for N iterations and converts results to TestRun records. Used by the main command handler.
Configuration API
The configuration module handles loading, parsing, and serializing the .ninety-nine.toml configuration file.
Loading
load_config
#![allow(unused)]
fn main() {
pub fn load_config(project_root: &Path) -> Result<Config, NinetyNineError>
}
Searches for .ninety-nine.toml in the given project root directory. Returns Config::default() if the file does not exist.
Errors: Returns ConfigParse if the file exists but contains invalid TOML, or ConfigIo if the file cannot be read.
default_config_toml
#![allow(unused)]
fn main() {
pub fn default_config_toml() -> Result<String, NinetyNineError>
}
Serializes Config::default() to pretty-printed TOML. Used by the init command.
backoff_base_delay
#![allow(unused)]
fn main() {
pub fn backoff_base_delay(strategy: &BackoffStrategy) -> Duration
}
Extracts the initial delay from a backoff strategy.
| Strategy | Base Delay |
|---|---|
None | 0ms |
Linear { delay_ms } | delay_ms |
Exponential { base_ms, .. } | base_ms |
Fibonacci { start_ms, .. } | start_ms |
Config Model
Config
Top-level configuration struct. All fields have sensible defaults.
#![allow(unused)]
fn main() {
pub struct Config {
pub detection: DetectionConfig,
pub retry: RetryConfig,
pub quarantine: QuarantineConfig,
pub storage: StorageConfig,
pub reporting: ReportingConfig,
}
}
DetectionConfig
Controls flakiness detection behavior.
#![allow(unused)]
fn main() {
pub struct DetectionConfig {
pub min_runs: u32,
pub confidence_threshold: f64,
pub window_size: u32,
pub parallel_runs: u32,
pub duration_regression: Option<DurationRegressionConfig>,
}
}
| Field | Default | Description |
|---|---|---|
min_runs | 10 | Minimum iterations per test |
confidence_threshold | 0.95 | Statistical confidence required to classify as flaky |
window_size | 100 | Maximum historical runs to consider |
parallel_runs | 3 | Number of concurrent test executions |
duration_regression | None | Duration regression tuning; None applies the defaults (enabled, 10-run history, 2 standard deviations) |
RetryConfig
Controls test retry behavior on failure.
#![allow(unused)]
fn main() {
pub struct RetryConfig {
pub unit_test_retries: u32,
pub backoff_strategy: BackoffStrategy,
pub max_retry_time_secs: u64,
}
}
| Field | Default | Description |
|---|---|---|
unit_test_retries | 2 | Maximum retries per failing test; every attempt is recorded as its own run |
backoff_strategy | Exponential(100ms, 2.0x, 5000ms) | Delay strategy between retries |
max_retry_time_secs | 300 | Time limit for a single test execution (seconds) |
BackoffStrategy
#![allow(unused)]
fn main() {
pub enum BackoffStrategy {
None,
Linear { delay_ms: u64 },
Exponential { base_ms: u64, factor: f64, max_ms: u64 },
Fibonacci { start_ms: u64, max_ms: u64 },
}
}
QuarantineConfig
Controls automatic and manual test quarantine.
#![allow(unused)]
fn main() {
pub struct QuarantineConfig {
pub enabled: bool,
pub auto_quarantine: bool,
pub threshold: QuarantineThreshold,
}
}
| Field | Default | Description |
|---|---|---|
enabled | true | Enable quarantine system |
auto_quarantine | false | Automatically quarantine tests exceeding thresholds |
QuarantineThreshold
#![allow(unused)]
fn main() {
pub struct QuarantineThreshold {
pub consecutive_failures: u32,
pub failure_rate: f64,
pub flakiness_score: f64,
}
}
| Field | Default | Description |
|---|---|---|
consecutive_failures | 3 | Consecutive failures before quarantine |
failure_rate | 0.20 | Failure rate threshold (20%) |
flakiness_score | 0.15 | Bayesian score threshold |
StorageConfig
#![allow(unused)]
fn main() {
pub struct StorageConfig {
pub backend: StorageBackendType,
pub retention_days: u32,
pub sqlite: Option<SqliteConfig>,
pub postgres: Option<PostgresConfig>,
}
}
| Field | Default | Description |
|---|---|---|
backend | Sqlite | Storage backend to use |
retention_days | 90 | Days to retain test run data |
StorageBackendType
#![allow(unused)]
fn main() {
pub enum StorageBackendType {
Sqlite,
Postgres,
}
}
SqliteConfig
#![allow(unused)]
fn main() {
pub struct SqliteConfig {
pub database_path: PathBuf,
}
}
Default path: $XDG_DATA_HOME/ninety-nine/ninety-nine.db
PostgresConfig
#![allow(unused)]
fn main() {
pub struct PostgresConfig {
pub connection_string: String,
pub pool_size: u32,
}
}
ReportingConfig
#![allow(unused)]
fn main() {
pub struct ReportingConfig {
pub console: ConsoleOutputConfig,
}
pub struct ConsoleOutputConfig {
pub summary_only: bool, // default: false
}
}
DurationRegressionConfig
#![allow(unused)]
fn main() {
pub struct DurationRegressionConfig {
pub enabled: bool,
pub min_history_runs: u32,
pub threshold: DurationThreshold,
}
pub enum DurationThreshold {
Multiplier(f64),
StdDev(f64),
}
}
See the Configuration Reference for TOML examples.
Error Types
All fallible operations in cargo-ninety-nine return Result<T, NinetyNineError>.
NinetyNineError
A comprehensive error enum covering all failure modes, derived with thiserror.
#![allow(unused)]
fn main() {
pub enum NinetyNineError {
ConfigParse { source: toml::de::Error },
ConfigIo { path: PathBuf, source: std::io::Error },
NoRunnerAvailable,
RunnerExecution { message: String },
InvalidConfig { message: String },
BinaryDiscovery { message: String },
TestListing { binary: PathBuf, message: String },
TestNotFound { name: String },
FilterParse { message: String },
Io { source: std::io::Error },
Json { source: serde_json::Error },
Storage { source: rusqlite::Error },
PostgresStorage { source: tokio_postgres::Error },
PostgresPool { message: String },
}
}
Error Categories
Configuration Errors
| Variant | Display Message | Cause |
|---|---|---|
ConfigParse | failed to parse config: {source} | TOML syntax error or invalid field value |
ConfigIo | failed to read config file {path}: {source} | File exists but cannot be read (permissions, etc.) |
InvalidConfig | invalid configuration: {message} | Logical configuration error (e.g., Postgres backend without connection config) |
Runner Errors
| Variant | Display Message | Cause |
|---|---|---|
NoRunnerAvailable | no test runner available: install cargo-nextest or use cargo test | Neither cargo-nextest nor cargo test found |
RunnerExecution | runner execution failed: {message} | Test binary failed to spawn or produced unexpected output |
BinaryDiscovery | binary discovery failed: {message} | cargo test --no-run failed or produced unparseable output |
TestListing | test listing failed for {binary}: {message} | Test binary --list invocation failed |
TestNotFound | test not found: {name} | Requested test does not exist in the project |
Filter Errors
| Variant | Display Message | Cause |
|---|---|---|
FilterParse | filter parse error: {message} | Invalid filter DSL syntax |
Storage Errors
| Variant | Display Message | Cause |
|---|---|---|
Storage | storage error: {source} | SQLite operation failed (auto-converted from rusqlite::Error) |
PostgresStorage | postgres storage error: {source} | PostgreSQL operation failed (auto-converted from tokio_postgres::Error) |
PostgresPool | postgres pool error: {message} | Connection pool exhausted or configuration error |
I/O and Serialization Errors
| Variant | Display Message | Cause |
|---|---|---|
Io | io error: {source} | Generic I/O failure (auto-converted from std::io::Error) |
Json | json serialization error: {source} | JSON serialization/deserialization failure (auto-converted from serde_json::Error) |
Automatic Conversions
The following From implementations allow ? propagation:
| Source Type | Target Variant |
|---|---|
std::io::Error | Io |
serde_json::Error | Json |
rusqlite::Error | Storage |
tokio_postgres::Error | PostgresStorage |
Error Handling in Practice
All errors are displayed to stderr in the main entry point and cause the process to exit with code 1:
#![allow(unused)]
fn main() {
// Simplified from main.rs
match run(args).await {
Ok(()) => {}
Err(e) => {
eprintln!("error: {e}");
std::process::exit(1);
}
}
}
When --verbose is enabled, the tracing subscriber is set to debug level, providing additional context before errors surface.
Architecture
Module Overview
cargo-ninety-nine
src/
main.rs Entry point, command dispatch
lib.rs Module re-exports
env.rs Git info, environment, CI provider detection
orchestrator.rs Test execution pipeline, session lifecycle, auto-quarantine
error.rs NinetyNineError enum (thiserror)
analysis/
mod.rs Shared failure_rate() helper
duration.rs Duration regression detection
pattern.rs Failure pattern detection (time-of-day, environmental)
trend.rs Trend calculation (improving/stable/degrading)
ci/
mod.rs CI re-exports
workflow.rs GitHub Actions / GitLab CI YAML generation
cli/
mod.rs CLI argument definitions (clap derive)
output.rs Console/JSON output formatting
export.rs JUnit XML, HTML, CSV, JSON export
config/
mod.rs Config loading from TOML
model.rs Config struct definitions with defaults
detector/
mod.rs Detector re-exports
bayesian.rs Bayesian flakiness scoring
filter/
mod.rs compile_filter(), build_eval_context()
ast.rs FilterExpr and Predicate AST nodes
lexer.rs Tokenizer for filter DSL
parser.rs Recursive descent parser
eval.rs Evaluator matching FilterExpr against TestMetadata
runner/
mod.rs NativeRunner, RunnerBackend, execute_iterations()
binary.rs Test binary discovery via cargo_metadata
listing.rs Test case enumeration from binaries
executor.rs Per-test execution with timeout/retry
detection.rs Runner availability check
storage/
mod.rs Storage trait (async), open_storage() factory
backend.rs StorageBackend enum with dispatch! macro
mapping.rs Shared row-to-domain-type conversion (RawTestRunRow, RawScoreRow)
sqlite.rs SQLite implementation (rusqlite, WAL mode)
postgres.rs PostgreSQL implementation (deadpool-postgres)
schema.rs SQL migration definitions
tui/
mod.rs TUI entry points, event loop, terminal guard, signal handlers
app.rs Application state (ScoresApp, HistoryApp, TableState, SortField)
input.rs Key event mapping to actions
render.rs Ratatui widget rendering (scores table, detail overlay, history)
types/
mod.rs Type re-exports
test_run.rs TestRun, TestOutcome, TestEnvironment
test_name.rs TestName newtype
flakiness.rs FlakinessScore, BayesianParams, FlakinessCategory
trend.rs TrendDirection, TrendSummary
session.rs RunSession, ActiveSession, QuarantineEntry
analysis.rs FailurePattern, PatternType
Test Execution Flow
CLI (clap)
|
v
Config Loader ---------> .ninety-nine.toml
|
v
Filter DSL (optional)
lexer --> parser --> FilterExpr AST
|
v
Runner Backend (NativeRunner)
+-------------------+
| Binary Discovery | cargo test --no-run --message-format json
| (cargo_metadata)|
+---------+---------+
|
v
+-------------------+
| Test Listing | binary --list --format terse
| (parallel via | (semaphore-bounded concurrency)
| tokio spawn) |
+---------+---------+
|
v
+-------------------+
| Filter Evaluation | FilterExpr evaluated against TestMetadata
| (if DSL given) | (loads flaky/quarantined sets from storage)
+---------+---------+
|
v
+-------------------+
| Executor | binary --exact test_name --nocapture
| (N iterations, | (parallel via spawn_blocking + Semaphore)
| timeout, retry) |
+-------------------+
|
| Vec<TestRun>
v
+-------------------+
| Bayesian Detector | Beta(alpha, beta) posterior
| + Analysis | --> FlakinessScore
| + Duration | --> DurationRegression (optional)
+-------------------+
|
+-----+-----+-------+
v v v
Storage Reporter TUI
(SQLite/ (console/ (ratatui,
Postgres) JSON/ interactive
export) scores/history)
Filter DSL Pipeline
The filter system has three stages:
-
Lexer (
filter/lexer.rs) – tokenizes the input string intoTokenvalues:LParen,RParen,And(&),Or(|),Not(!), andIdent(String). -
Parser (
filter/parser.rs) – recursive descent parser that builds aFilterExprAST. Supports binary operators (&,|), unary negation (!), parenthesized grouping, and function-call predicates liketest(pattern),package(name),binary(name),kind(lib|bin|test|example). Bare identifiers are resolved as keywords (flaky,quarantined,all) or treated as test name regex patterns. -
Evaluator (
filter/eval.rs) – evaluates aFilterExpragainst aTestMetadatastruct containing the test name, package name, binary name, and binary kind. AnEvalContextholds pre-loaded sets of flaky and quarantined test names from storage.
Storage Abstraction
The Storage trait defines 13 async methods for persisting and querying test data. Two backends implement it:
-
SqliteStorage – uses
rusqlitewith WAL mode andMutex<Connection>for thread safety. Synchronous operations run inside async method signatures. Migrations usePRAGMA user_version. -
PostgresStorage – uses
deadpool-postgresfor connection pooling with configurable pool size and timeouts. Migrations use aschema_migrationstable.
The StorageBackend enum wraps both backends and dispatches calls via a dispatch! macro, keeping the orchestration layer backend-agnostic.
#![allow(unused)]
fn main() {
pub enum StorageBackend {
Sqlite(SqliteStorage),
Postgres(PostgresStorage),
}
}
The open_storage() factory function reads the config to determine which backend to initialize.
Type System
TestName
A newtype wrapper around String that prevents confusion with other string fields (branch names, commit hashes, error messages). Implements Deref<Target=str>, AsRef<str>, Display, From<String>, From<&str>, and PartialEq<str>. Used in TestRun, FlakinessScore, QuarantineEntry, TrendSummary, DurationRegression, and TestCase.
ActiveSession
Session lifecycle type. ActiveSession::start() creates a running session, and to_run_session() produces the storable RunSession snapshot without consuming it; the orchestrator finishes the session by id once the suite completes.
FlakinessScore and FlakinessCategory
FlakinessScore holds the Bayesian posterior parameters (alpha, beta, posterior mean/variance, credible interval) alongside aggregate statistics (pass rate, fail rate, consecutive failures). FlakinessCategory classifies scores into five levels:
| Score Range | Category |
|---|---|
| < 0.01 | Stable |
| 0.01 - 0.05 | Occasional |
| 0.05 - 0.15 | Moderate |
| 0.15 - 0.30 | Frequent |
| >= 0.30 | Critical |
Key Design Decisions
Native Test Runner
Instead of wrapping cargo test as a subprocess for the full run, the tool uses a three-layer native pipeline:
- Binary discovery – uses
cargo_metadatato parsecargo test --no-runJSON output, extracting test binary paths with package name and binary kind. - Test listing – executes each binary with
--list --format terseto enumerate individual test names. Runs in parallel across binaries usingtokio::spawn_blockingwith semaphore-bounded concurrency. - Per-test execution – runs each test individually via
duct, using--exactfor isolation. Supports configurable timeout, retry count, and retry delay.
This gives precise per-test timing, retry control, and outcome classification without parsing human-readable test output.
Parallel Execution
Test execution uses tokio::spawn_blocking with a Semaphore to bound concurrency. The semaphore limit comes from detection.parallel_runs in the config (default: 3). Each test iteration acquires a permit before spawning a blocking task.
Subprocess Management
Test execution uses duct for subprocess control with:
- Poll-based timeout –
try_wait()in a loop with a 50ms polling interval,kill()on deadline - Output capture – stdout and stderr captured for failure analysis
- Unchecked mode – non-zero exit codes are handled as test outcomes, not errors
Storage Backend Dispatch
The StorageBackend enum uses a dispatch! macro to forward all 13 trait methods to the underlying backend. This avoids 130+ lines of boilerplate match arms while keeping the dispatch zero-cost.
Error Handling
All errors flow through NinetyNineError (thiserror), with variants for each subsystem: config, storage, binary discovery, test listing, runner execution, filter parsing, Postgres pool, and I/O. The ? operator propagates errors to the top-level handler.
Bayesian Detection
Overview
cargo ninety-nine uses Bayesian inference to estimate the probability that a test is flaky. This approach provides calibrated uncertainty estimates rather than simple pass/fail ratios.
The Model
Prior
The model starts with a uniform (uninformative) prior: Beta(1, 1). This represents no prior knowledge — any flakiness probability from 0 to 1 is equally likely before observing data.
Posterior Update
After observing test runs, the posterior distribution is:
Beta(alpha, beta)
where:
alpha = prior_alpha + failures
beta = prior_beta + passes
The posterior mean is used as the flakiness probability:
P(flaky) = alpha / (alpha + beta)
Credible Interval
A 95% credible interval is computed from the Beta distribution’s inverse CDF:
CI = [Beta.inverse_cdf(0.025), Beta.inverse_cdf(0.975)]
This interval narrows as more data is collected.
Confidence
Confidence is derived from the width of the credible interval:
confidence = 1.0 - (CI_upper - CI_lower)
Narrow intervals yield high confidence; wide intervals yield low confidence.
Classification
A test is classified as flaky when:
P(flaky) > 0.01 AND confidence >= confidence_threshold
The confidence_threshold is configurable (default: 0.95).
Flakiness Categories
| Category | P(flaky) Range |
|---|---|
| Stable | < 1% |
| Occasional | 1% — 5% |
| Moderate | 5% — 15% |
| Frequent | 15% — 30% |
| Critical | > 30% |
Practical Implications
Number of Runs
With the Beta(1,1) prior:
- 10 runs, 1 failure: P(flaky) = 2/12 = 16.7%, wide CI → low confidence
- 100 runs, 1 failure: P(flaky) = 2/102 = 2.0%, narrow CI → high confidence
- 100 runs, 0 failures: P(flaky) = 1/102 = 1.0%, classified as Stable
More runs yield narrower credible intervals and more reliable classifications.
Prior Effect
The uniform prior has minimal effect when there are many observations. With 100+ runs, the prior contributes < 2% to the posterior. For small sample sizes (< 10 runs), the prior pulls estimates toward 50%, which is conservative.
Stored Parameters
Each FlakinessScore record includes the full Bayesian parameters:
| Field | Description |
|---|---|
alpha | Posterior alpha (prior + failures) |
beta | Posterior beta (prior + passes) |
posterior_mean | alpha / (alpha + beta) |
posterior_variance | (alpha * beta) / (total^2 * (total + 1)) |
credible_interval_lower | 2.5th percentile of posterior |
credible_interval_upper | 97.5th percentile of posterior |
Analysis and Patterns
Pattern Detection
After running tests, cargo ninety-nine analyzes failure data to identify patterns that might explain why tests are flaky.
Pattern Types
| Pattern | Detection Method |
|---|---|
| TimeOfDay | Failures concentrated at a specific hour (3x expected rate) |
| Environmental | Failure rate differs > 15% between CI and local environments |
| Random | Failures present but no discernible pattern (fallback) |
Time-of-Day Detection
Requires at least 5 failures. Counts failures by hour (0-23 UTC) and computes:
concentration = max_hour_count / expected_per_hour
If concentration >= 3.0, a TimeOfDay pattern is reported with:
correlation:min(concentration - 1.0, 1.0)examples: the peak hour identified
This pattern suggests timing-dependent failures (e.g., midnight log rotation, scheduled background jobs).
Environmental Detection
Requires at least 3 CI runs and 3 local runs. Compares failure rates between environments:
diff = |ci_fail_rate - local_fail_rate|
If diff >= 0.15, an Environmental pattern is reported, identifying which environment has the higher failure rate. This pattern suggests resource-dependent failures (e.g., CPU count, memory limits, file system speed).
Random Pattern
If failures exist but no specific pattern is detected, a Random pattern is reported. This is the fallback – it means failures occur but do not correlate with time or environment.
Trend Analysis
Trend analysis compares recent failure rates to previous failure rates within a sliding window.
How It Works
- Filter runs for a specific test name
- Take up to
window_sizemost recent runs (configurable, default: 100) - Split into two halves: recent (first half) and previous (second half)
- Compute failure rate for each half
- Classify the delta:
| Delta | Direction |
|---|---|
| > 5% | Degrading |
| < -5% | Improving |
| within +/- 5% | Stable |
Requirements
At least 4 runs are required for trend calculation. Fewer runs return None.
Where Trends Are Shown
cargo ninety-nine status <test_name>– shows trend direction and delta- After
test– degrading trends are highlighted in the post-detection analysis
Duration Regression Analysis
Duration regression detection identifies tests whose execution time has significantly increased compared to their historical average. This helps catch performance regressions early, even when tests still pass.
How It Works
- Collect all run durations for a test (most recent first)
- Require at least
min_history_runstotal data points - Separate the latest run from historical runs
- Compute the mean and sample standard deviation of historical durations only (the latest run is excluded to avoid polluting the baseline with the spike being measured)
- Apply a standard deviation floor of 1% of the mean (handles identical-duration histories where raw std_dev is zero)
- Compute the z-score:
deviation = (latest_duration - historical_mean) / effective_std_dev
- If
deviation > threshold, a regression is reported with the actual slowdown ratio (latest / mean)
Configuration
Duration regression detection runs by default with a 10-run history requirement and a 2 standard-deviation threshold. Provide the section in .ninety-nine.toml to tune those values, or set enabled = false to switch the check off:
[detection.duration_regression]
enabled = true
min_history_runs = 10
# Flag when latest run exceeds 2.5 standard deviations above the mean
[detection.duration_regression.threshold]
StdDev = 2.5
Two threshold variants are available:
| Variant | Description | Example |
|---|---|---|
StdDev(f64) | Number of standard deviations above the mean | StdDev = 2.0 flags tests running >2 sigma slower |
Multiplier(f64) | Multiple of the historical mean | Multiplier = 3.0 flags tests taking >3x their average |
Output
When a regression is detected, the CLI displays:
SLOW tests::my_test — 500ms (mean: 100ms, 5.0x)
The DurationRegression struct contains:
| Field | Description |
|---|---|
test_name | Fully qualified name of the affected test |
current_ms | Duration of the latest run in milliseconds |
mean_ms | Historical mean duration in milliseconds (excludes latest) |
std_dev_ms | Standard deviation of historical durations |
deviation_factor | How many standard deviations above the historical mean |
The display shows the actual slowdown ratio (current_ms / mean_ms) rather than the z-score.
Edge Cases
- If the effective standard deviation is near zero after flooring, no regression is reported
- If fewer than
min_history_runstotal runs exist, no regression check is performed - If only one historical run exists (after separating the latest), the mean is that single value and the floor applies
Storage
Storage Trait
The Storage trait defines the async interface for all data persistence. Both backends implement all 13 methods:
#![allow(unused)]
fn main() {
pub trait Storage: Send + Sync {
async fn store_session(&self, session: &RunSession) -> Result<(), NinetyNineError>;
async fn finish_session(&self, session_id: &Uuid, test_count: u32, flaky_count: u32) -> Result<(), NinetyNineError>;
async fn store_test_run(&self, run: &TestRun, session_id: &Uuid) -> Result<(), NinetyNineError>;
async fn store_flakiness_score(&self, score: &FlakinessScore) -> Result<(), NinetyNineError>;
async fn get_test_runs(&self, test_name: &str, limit: u32) -> Result<Vec<TestRun>, NinetyNineError>;
async fn get_recent_sessions(&self, limit: u32) -> Result<Vec<RunSession>, NinetyNineError>;
async fn get_all_scores(&self) -> Result<Vec<FlakinessScore>, NinetyNineError>;
async fn get_score(&self, test_name: &str) -> Result<Option<FlakinessScore>, NinetyNineError>;
async fn quarantine_test(&self, test_name: &str, reason: &str, score: f64, auto: bool) -> Result<(), NinetyNineError>;
async fn unquarantine_test(&self, test_name: &str) -> Result<(), NinetyNineError>;
async fn get_quarantined_tests(&self) -> Result<Vec<QuarantineEntry>, NinetyNineError>;
async fn is_quarantined(&self, test_name: &str) -> Result<bool, NinetyNineError>;
async fn purge_older_than(&self, days: u32) -> Result<u64, NinetyNineError>;
}
}
The StorageBackend enum wraps both backends and dispatches calls via a dispatch! macro:
#![allow(unused)]
fn main() {
pub enum StorageBackend {
Sqlite(SqliteStorage),
Postgres(PostgresStorage),
}
}
The open_storage() factory function reads the config to initialize the correct backend.
SQLite Backend (Default)
SQLite is the default backend, using rusqlite with the bundled SQLite library. No external database is required.
Default Location
$XDG_DATA_HOME/ninety-nine/ninety-nine.db
On Linux this is typically ~/.local/share/ninety-nine/ninety-nine.db. The parent directory is created automatically if it does not exist.
Features
- WAL mode – enabled on open for concurrent read access during detection
- Foreign keys – enforced via
PRAGMA foreign_keys=ON - Thread safety –
Mutex<Connection>guards all access; async trait methods run synchronously within async signatures - Bundled SQLite – no system SQLite dependency required
Configuration
[storage]
backend = "Sqlite"
retention_days = 90
[storage.sqlite]
database_path = "/custom/path/ninety-nine.db"
If storage.sqlite is omitted, the default path is used.
Migrations
Schema migrations use SQLite’s PRAGMA user_version. Each migration increments the version. Migrations are idempotent – running the tool against an already-migrated database is safe.
PostgreSQL Backend
PostgreSQL support uses deadpool-postgres for connection pooling with tokio-postgres for async queries.
Features
- Connection pooling – configurable pool size via
deadpool-postgres - Pool timeouts – 30s wait, 10s create, 5s recycle
- Native async – all queries use async
tokio-postgresdirectly - Schema migrations – tracked via a
schema_migrationstable with version and timestamp
Configuration
[storage]
backend = "Postgres"
retention_days = 90
[storage.postgres]
connection_string = "postgresql://user:password@localhost:5432/ninety_nine"
pool_size = 8
Note: Selecting
backend = "Postgres"without providing[storage.postgres]will result in a configuration error at startup.
Migrations
PostgreSQL migrations use a schema_migrations table instead of pragmas. The migration system tracks applied versions and only runs new migrations. The schema is identical to SQLite in structure.
Schema
The database has four tables, identical in structure across both backends.
run_sessions
Tracks each detection run session.
| Column | Type | Description |
|---|---|---|
id | TEXT (UUID) | Session identifier |
started_at | TEXT/TIMESTAMPTZ | When the session started |
finished_at | TEXT/TIMESTAMPTZ | When the session finished (nullable) |
test_count | INTEGER | Total tests analyzed |
flaky_count | INTEGER | Tests classified as flaky |
commit_hash | TEXT | Git commit at time of run |
branch | TEXT | Git branch at time of run |
test_runs
Individual test execution results.
| Column | Type | Description |
|---|---|---|
id | TEXT (UUID) | Run identifier |
session_id | TEXT (FK) | Parent session |
test_name | TEXT | Fully qualified test name |
test_path | TEXT | Binary path |
outcome | TEXT | passed, failed, timeout, panic, ignored |
duration_ms | INTEGER/BIGINT | Execution time in milliseconds |
timestamp | TEXT/TIMESTAMPTZ | When this run occurred |
commit_hash | TEXT | Git commit |
branch | TEXT | Git branch |
retry_count | INTEGER | Number of retries used |
error_message | TEXT | Stderr/stdout on failure (nullable) |
stack_trace | TEXT | Stack trace if available (nullable) |
env_os | TEXT | Operating system |
env_rust_version | TEXT | Rust toolchain version |
env_cpu_count | INTEGER | CPU core count |
env_memory_gb | REAL/DOUBLE PRECISION | System memory in GB |
env_is_ci | INTEGER/BOOLEAN | Whether running in CI |
env_ci_provider | TEXT | CI provider name (nullable) |
Indexed on test_name, timestamp, and session_id.
flakiness_scores
Latest computed flakiness scores, upserted on test_name.
| Column | Type | Description |
|---|---|---|
test_name | TEXT (PK) | Fully qualified test name |
probability_flaky | REAL | Bayesian P(flaky) |
confidence | REAL | 1 - credible interval width |
pass_rate | REAL | Passes / total |
fail_rate | REAL | Failures / total |
total_runs | INTEGER | Number of runs |
consecutive_failures | INTEGER | Trailing failures |
last_updated | TEXT/TIMESTAMPTZ | Last computation time |
alpha | REAL | Beta distribution alpha parameter |
beta | REAL | Beta distribution beta parameter |
posterior_mean | REAL | Posterior mean |
posterior_variance | REAL | Posterior variance |
ci_lower | REAL | 95% credible interval lower bound |
ci_upper | REAL | 95% credible interval upper bound |
quarantine
Quarantined test records.
| Column | Type | Description |
|---|---|---|
test_name | TEXT (PK) | Fully qualified test name |
quarantined_at | TEXT/TIMESTAMPTZ | When quarantined |
reason | TEXT | Reason for quarantine |
flakiness_score | REAL | P(flaky) at time of quarantine |
auto_quarantined | INTEGER/BOOLEAN | Whether auto-quarantined |
Indexed on quarantined_at.
Data Retention
Old test runs are automatically purged based on the retention_days config (default: 90 days). Purging happens after each test session. Only the test_runs table is purged – flakiness_scores and quarantine entries persist.
[storage]
retention_days = 30
CLI Reference
Synopsis
cargo ninety-nine [OPTIONS] <COMMAND>
Global Options
| Option | Default | Description |
|---|---|---|
--project-dir <PATH> | . | Project root directory |
--output <FORMAT> | console | Output format: console or json |
-N, --non-interactive | false | Disable TUI, use plain text output |
-v, --verbose | false | Enable verbose output |
By default, test, status, and history launch an interactive TUI. Use -N for CI pipelines, scripts, or piped output.
Commands
diagnose
Multi-phase stress → isolation → classify pipeline. See Multi-phase Diagnose.
cargo ninety-nine diagnose [OPTIONS] [FILTER_EXPR]
| Argument | Default | Description |
|---|---|---|
[FILTER_EXPR] | none | Filter expression (same DSL as test) |
--stress-runs <N> | config (3) | Full-binary multi-threaded suite runs |
--isolation-runs <N> | config (10) | Serial --exact runs per candidate |
--record | false | Attempt rr recording for Intrinsic failures |
--no-record | false | Disable recording even if config enables it |
--record-dir <PATH> | config | Directory for rr traces |
--confidence <F> | config | Bayesian confidence threshold |
test
Run tests and detect flakiness. Each discovered test is executed multiple times, scored with Bayesian inference, and results are stored for trend analysis.
cargo ninety-nine test [OPTIONS] [FILTER_EXPR]
| Argument/Option | Default | Description |
|---|---|---|
[FILTER_EXPR] | none | Filter expression (DSL or test name pattern) |
-n, --iterations <N> | from config (min_runs, default 10) | Number of times to run each test |
--confidence <FLOAT> | from config (confidence_threshold, default 0.95) | Confidence threshold for flaky classification |
Examples:
# Run all tests 10 times each (defaults)
cargo ninety-nine test
# Run tests matching a substring
cargo ninety-nine test my_module
# Run with more iterations and stricter confidence
cargo ninety-nine test -n 25 --confidence 0.99
# Use filter DSL to run only flaky tests
cargo ninety-nine test "flaky"
# Combine filter predicates
cargo ninety-nine test "test(my_module) & !quarantined"
Filter DSL
The optional FILTER_EXPR argument accepts a domain-specific language for filtering tests. If the expression contains no DSL operators, it is treated as a plain test name substring filter.
Predicates:
| Predicate | Description |
|---|---|
test(pattern) | Match test names by regex pattern |
package(name) | Match tests in a package (substring match) |
binary(name) | Match tests from a specific binary (substring match) |
kind(type) | Match by binary kind: lib, bin, test, example |
flaky | Match tests previously detected as flaky |
quarantined | Match quarantined tests |
all | Match all tests |
bare_word | Treated as test(bare_word) regex pattern |
Operators:
| Operator | Meaning |
|---|---|
& | AND – both sides must match |
| | OR – either side must match |
! | NOT – negate the following expression |
( ) | Grouping |
Examples:
# Tests matching a regex
cargo ninety-nine test "test(my_module::.*)"
# Flaky tests that are not quarantined
cargo ninety-nine test "flaky & !quarantined"
# Tests in a specific package or binary kind
cargo ninety-nine test "package(my_crate) | kind(test)"
# Complex expression with grouping
cargo ninety-nine test "(flaky | test(slow)) & !quarantined"
init
Initialize a .ninety-nine.toml configuration file in the project root.
cargo ninety-nine init [OPTIONS]
| Option | Description |
|---|---|
--force | Overwrite existing config file |
status
Show flakiness status for tests. Without a test name, launches the interactive TUI (or prints a table with -N).
cargo ninety-nine status [TEST_NAME]
| Argument | Description |
|---|---|
[TEST_NAME] | Show detailed status for a specific test. Omit to browse all scores interactively. |
When a test name is provided, shows: category, P(flaky), pass rate, total runs, consecutive failures, credible interval, trend direction, failure patterns, and recent run history.
history
Show detection session history. Launches the interactive TUI by default (or prints a table with -N).
cargo ninety-nine history [OPTIONS] [FILTER]
| Argument/Option | Default | Description |
|---|---|---|
[FILTER] | none | Filter by test name |
-n, --limit <N> | 20 | Maximum sessions to show |
export
Export flakiness data to a file.
cargo ninety-nine export <FORMAT> <PATH>
| Argument | Values | Description |
|---|---|---|
<FORMAT> | junit, html, csv, json | Export format |
<PATH> | file path | Output file path |
quarantine
Manage test quarantine.
quarantine list
cargo ninety-nine quarantine list
Lists all quarantined tests with their quarantine date, reason, flakiness score, and whether they were auto-quarantined.
quarantine add
cargo ninety-nine quarantine add <TEST_NAME> [OPTIONS]
| Option | Default | Description |
|---|---|---|
--reason <TEXT> | "manually quarantined" | Reason for quarantine |
quarantine remove
cargo ninety-nine quarantine remove <TEST_NAME>
ci
CI integration helpers.
ci generate
Generate a CI workflow file for flaky test detection.
cargo ninety-nine ci generate <PROVIDER> [PATH]
| Argument | Values | Description |
|---|---|---|
<PROVIDER> | github, gitlab | CI provider |
[PATH] | file path | Output file path (default: stdout) |
Configuration Reference
All configuration lives in .ninety-nine.toml at the project root. Every field has a default value – you only need to specify overrides.
[detection]
Controls how tests are discovered and analyzed.
| Field | Type | Default | Description |
|---|---|---|---|
min_runs | u32 | 10 | Minimum number of runs per test |
confidence_threshold | f64 | 0.95 | Confidence level required to classify as flaky |
window_size | u32 | 100 | Number of recent runs used for trend analysis |
parallel_runs | u32 | 3 | Number of concurrent test executions |
Detection uses Beta distribution posterior inference with conjugate priors. A test is only ever classified as flaky once at least one failure has actually been observed; long histories of clean passes always classify as stable.
[detection.duration_regression]
Optional configuration for detecting tests whose execution time has significantly increased.
| Field | Type | Default | Description |
|---|---|---|---|
enabled | bool | – | Enable duration regression detection |
min_history_runs | u32 | – | Minimum historical runs required before checking |
threshold | enum | – | How to determine a regression (see below) |
When the section is omitted entirely, the check runs with its long-standing defaults: a 10-run history requirement and a threshold of 2 standard deviations. Provide the section to tune those values, or set enabled = false to switch the check off:
[detection.duration_regression]
enabled = true
min_history_runs = 10
# Option A: flag when latest duration exceeds N standard deviations above the mean
[detection.duration_regression.threshold]
StdDev = 2.0
# Option B: flag when latest duration exceeds mean * multiplier
# [detection.duration_regression.threshold]
# Multiplier = 3.0
Threshold variants:
| Variant | Description |
|---|---|
StdDev(f64) | Flag when the latest duration is more than N standard deviations above the historical mean |
Multiplier(f64) | Flag when the latest duration exceeds the historical mean multiplied by this factor |
[diagnose]
Controls the multi-phase diagnose command (stress / isolation / optional rr).
| Field | Type | Default | Description |
|---|---|---|---|
stress_runs | u32 | 3 | Full-binary multi-threaded suite iterations |
isolation_runs | u32 | 10 | Serial isolation iterations per candidate |
stress_threads | u32 | 0 | libtest threads (0 = host parallelism) |
stress_timeout_secs | u64 | 300 | Timeout per stress binary iteration |
record | bool | false | Attempt rr recording for Intrinsic failures |
record_dir | path | .ninety-nine/recordings | rr output directory |
record_attempts | u32 | 10 | Max rr attempts per Intrinsic candidate |
chaos | bool | false | Pass --chaos to rr when recording |
[detection] multi-phase
| Field | Type | Default | Description |
|---|---|---|---|
multi_phase | bool | false | When true, test runs diagnose for candidates then Bayesian suite for the rest |
[quarantine.by_class]
| Field | Type | Default | Description |
|---|---|---|---|
intrinsic | bool | true | Auto-quarantine Intrinsic diagnose classes |
contention | bool | false | Auto-quarantine Contention classes |
broken | bool | true | Auto-quarantine Broken classes |
[retry]
Controls retry behavior when tests fail.
| Field | Type | Default | Description |
|---|---|---|---|
unit_test_retries | u32 | 2 | Number of retries for failed tests |
backoff_strategy | enum | Exponential | Backoff strategy between retries |
max_retry_time_secs | u64 | 300 | Time limit for a single test execution (seconds) |
Every attempt is recorded as its own run, so a test that fails and then passes on retry contributes both outcomes to its flakiness score — recovery on retry is itself flaky evidence, and raising the retry count therefore gathers more evidence per iteration rather than hiding failures.
Backoff Strategies
# No delay between retries
[retry]
backoff_strategy = "None"
# Fixed delay
[retry.backoff_strategy]
Linear = { delay_ms = 200 }
# Exponential backoff (default: base_ms=100, factor=2.0, max_ms=5000)
[retry.backoff_strategy]
Exponential = { base_ms = 100, factor = 2.0, max_ms = 5000 }
# Fibonacci sequence delays
[retry.backoff_strategy]
Fibonacci = { start_ms = 100, max_ms = 5000 }
[quarantine]
Controls test quarantine behavior.
| Field | Type | Default | Description |
|---|---|---|---|
enabled | bool | true | Enable quarantine functionality |
auto_quarantine | bool | false | Automatically quarantine tests exceeding thresholds |
[quarantine.threshold]
Thresholds for auto-quarantine. A test must be classified as flaky AND exceed at least one threshold.
| Field | Type | Default | Description |
|---|---|---|---|
consecutive_failures | u32 | 3 | Consecutive trailing failures |
failure_rate | f64 | 0.20 | Overall failure rate |
flakiness_score | f64 | 0.15 | Bayesian P(flaky) |
[storage]
Controls data persistence. Two backends are supported: SQLite (default) and PostgreSQL.
| Field | Type | Default | Description |
|---|---|---|---|
backend | enum | "Sqlite" | Storage backend: Sqlite or Postgres |
retention_days | u32 | 90 | Days to keep test run data; sessions left without any runs by the purge are removed with them |
[storage.sqlite]
SQLite-specific settings. Only used when backend = "Sqlite".
| Field | Type | Default | Description |
|---|---|---|---|
database_path | path | $XDG_DATA_HOME/ninety-nine/ninety-nine.db | Path to the SQLite database file |
[storage]
backend = "Sqlite"
retention_days = 90
[storage.sqlite]
database_path = ".ninety-nine/data.db"
[storage.postgres]
PostgreSQL-specific settings. Required when backend = "Postgres".
| Field | Type | Default | Description |
|---|---|---|---|
connection_string | string | – | PostgreSQL connection URL |
pool_size | u32 | – | Connection pool size |
[storage]
backend = "Postgres"
retention_days = 90
[storage.postgres]
connection_string = "postgresql://user:password@localhost:5432/ninety_nine"
pool_size = 8
Warning: Setting
backend = "Postgres"without providing[storage.postgres]will cause a startup error.
[reporting]
Controls output behavior.
[reporting.console]
| Field | Type | Default | Description |
|---|---|---|---|
summary_only | bool | false | Show only summary counts instead of full table |
Full Example
[detection]
min_runs = 20
confidence_threshold = 0.99
window_size = 50
parallel_runs = 4
[detection.duration_regression]
enabled = true
min_history_runs = 10
[detection.duration_regression.threshold]
StdDev = 2.5
[retry]
unit_test_retries = 3
max_retry_time_secs = 120
[retry.backoff_strategy]
Exponential = { base_ms = 200, factor = 2.0, max_ms = 10000 }
[quarantine]
enabled = true
auto_quarantine = true
[quarantine.threshold]
consecutive_failures = 5
failure_rate = 0.15
flakiness_score = 0.10
[storage]
backend = "Sqlite"
retention_days = 60
[storage.sqlite]
database_path = ".ninety-nine/data.db"
[reporting.console]
summary_only = false