Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

cargo ninety-nine

cargo ninety-nine logo

A cargo plugin for finding and tracking flaky tests in Rust projects using Bayesian inference.

Flaky tests – tests that sometimes pass and sometimes fail without code changes – erode trust in your test suite, waste CI time, and mask real regressions. cargo ninety-nine runs each test multiple times, applies Bayesian statistical analysis to compute a flakiness probability, and tracks results over time in a local SQLite database.

Key Features

  • Bayesian flakiness detection – computes posterior probability of flakiness using Beta distribution, not just pass/fail ratios
  • Filter DSL – expressive query language to target tests by name, package, binary, kind, or flakiness status
  • Pattern analysis – detects time-of-day and environmental (CI vs local) failure patterns
  • Trend tracking – monitors whether tests are improving, stable, or degrading over time
  • Interactive TUI – terminal interface for browsing scores with sorting, filtering, and detail drill-down
  • Quarantine management – manually or automatically quarantine flaky tests that exceed thresholds
  • Multiple export formats – JUnit XML, HTML reports, CSV, and JSON for integration with other tools
  • CI workflow generation – generates ready-to-use GitHub Actions and GitLab CI configurations
  • Persistent storage – SQLite database with WAL mode for concurrent access, automatic data retention

Quick Example

# Initialize configuration
cargo ninety-nine init

# Run flaky test detection (10 iterations per test)
cargo ninety-nine test -n 10

# Run only flaky, non-quarantined tests
cargo ninety-nine test "flaky & !quarantined"

# Check a specific test's history
cargo ninety-nine status tests::my_flaky_test

# Export results as JSON
cargo ninety-nine export json results.json

How It Works

  1. Discover – compiles test binaries and lists all test cases
  2. Filter – applies filter expressions to select which tests to run
  3. Execute – runs each test N times with configurable concurrency and timeouts
  4. Analyze – applies Bayesian inference to compute P(flaky) with credible intervals
  5. Report – displays results with category labels (Stable, Occasional, Moderate, Frequent, Critical)
  6. Store – persists scores and run history for trend analysis across sessions

Getting Started

Prerequisites

  • Rust 1.85+ (Edition 2024)
  • cargo (included with Rust)
  • Optionally, cargo-nextest for enhanced test running

Installation

Install from crates.io:

cargo install cargo-ninety-nine

Or build from source:

git clone https://github.com/glottologist/ninety-nine.git
cd ninety-nine
cargo install --path .

Verify Installation

cargo ninety-nine --version

Initialize a Project

Navigate to your Rust project and create a configuration file:

cd /path/to/your/project
cargo ninety-nine init

This creates a .ninety-nine.toml file with sensible defaults. You can overwrite an existing config with:

cargo ninety-nine init --force

First Test Run

Run detection with default settings (10 iterations per test):

cargo ninety-nine test

Filter to specific tests using a name pattern:

cargo ninety-nine test "my_module::tests"

Use the filter DSL to target specific subsets:

cargo ninety-nine test "flaky & !quarantined"

Increase iterations for higher confidence:

cargo ninety-nine test -n 50 --confidence 0.99

Output Formats

Add --output json to any command for machine-readable output:

cargo ninety-nine status --output json
cargo ninety-nine history --output json

Configuration

Configuration is stored in .ninety-nine.toml at the project root. All fields are optional — missing fields use defaults.

Generating a Config File

cargo ninety-nine init

Minimal Configuration

An empty .ninety-nine.toml file uses all defaults. You only need to specify values you want to change:

[detection]
min_runs = 20
confidence_threshold = 0.99

Full Configuration Reference

See Configuration Reference for all available options and their defaults.

Config Loading

The config file is loaded from the --project-dir path (defaults to .). If no config file exists, all default values are used. Unknown fields in the TOML file are silently ignored, so the config file is forward-compatible with newer versions.

Usage Overview

cargo ninety-nine provides seven subcommands:

CommandPurpose
testRun tests repeatedly and compute flakiness scores
initCreate a default configuration file
statusView current flakiness scores and test detail
historyView past detection sessions
exportExport results to JUnit XML, HTML, CSV, or JSON
quarantineManage test quarantine (list, add, remove)
ciGenerate CI workflow files

Interactive Mode

By default, test, status, and history launch an interactive TUI with sortable tables, category filtering, and detail drill-down. Pass --non-interactive (or -N) to use plain text output instead – useful for CI, scripts, or piped output.

Global Options

These options apply to all subcommands:

--project-dir <PATH>       Project root directory (default: .)
--output <FORMAT>          Output format: console or json (default: console)
--non-interactive, -N      Disable TUI, use plain text output
--verbose, -v              Verbose output during detection

Typical Workflow

  1. Initialize – cargo ninety-nine init
  2. Test – cargo ninety-nine test -n 20
  3. Filter – cargo ninety-nine test "flaky & !quarantined" (see Filter DSL)
  4. Investigate – cargo ninety-nine status tests::suspect_test
  5. Quarantine – cargo ninety-nine quarantine add tests::flaky_test --reason "timing-dependent"
  6. Export – cargo ninety-nine export html report.html
  7. Automate – cargo ninety-nine ci generate github

Running Tests

The test subcommand is the core functionality – it discovers, runs, and analyzes tests for flakiness.

Basic Usage

# Run all tests 10 times each (default)
cargo ninety-nine test

# Filter to specific tests by name pattern
cargo ninety-nine test "my_module::tests"

# 50 iterations with 99% confidence threshold
cargo ninety-nine test -n 50 --confidence 0.99

Options

OptionDefaultDescription
<filter_expr>noneFilter expression – a test name regex or DSL expression
-n, --iterations10Number of times to run each test
--confidence0.95Confidence threshold for classifying a test as flaky

Filter Expressions

The positional argument accepts either a plain regex pattern or a full filter DSL expression:

# Regex pattern matching test names
cargo ninety-nine test "my_module::tests"

# DSL: only tests marked flaky that are not quarantined
cargo ninety-nine test "flaky & !quarantined"

# DSL: tests in a specific package
cargo ninety-nine test "package(my_crate)"

# DSL: combine predicates
cargo ninety-nine test "test(network) & kind(test) & !quarantined"

See the Filter DSL guide for the full syntax reference.

What Happens During a Run

  1. Binary discovery – runs cargo test --no-run --message-format json-render-diagnostics to compile and locate test binaries
  2. Test listing – executes each binary with --list --format terse to enumerate test cases
  3. Filtering – evaluates the filter expression (if provided) against each test’s metadata
  4. Parallel execution – runs each matched test individually N times using --exact --nocapture flags
  5. Bayesian scoring – computes P(flaky) using Beta distribution with uniform prior
  6. Storage – writes results to SQLite database for trend tracking
  7. Reporting – displays results sorted by flakiness probability

Understanding the Output

Console Report

Flaky Test Detection Report

Test                                               Runs     Pass%   P(flaky)     Category
--------------------------------------------------------------------------------------------
tests::race_condition                                10     60.0%      33.3%     Frequent
tests::timing_dependent                              10     80.0%      16.7%     Moderate
tests::stable_test                                   10    100.0%       1.0%       Stable

Flakiness Categories

CategoryP(flaky) RangeMeaning
Stable< 1%Reliably passes
Occasional1% - 5%Rare failures, may be acceptable
Moderate5% - 15%Notable flakiness, worth investigating
Frequent15% - 30%Significant problem
Critical> 30%Severely broken, likely a bug

Post-Run Analysis

After the main report, the tool displays:

  • Detected Patterns – time-of-day or environmental correlations in failure data
  • Degrading Trends – tests whose failure rate has increased between recent and previous runs

Auto-Quarantine

If quarantine.auto_quarantine = true in config, tests exceeding the configured thresholds are automatically quarantined after detection. See Quarantine Management.

Verbose Mode

Use -v for per-test progress output instead of the progress bar:

cargo ninety-nine test -v

Summary Mode

Set reporting.console.summary_only = true in config to show only aggregate counts:

Summary: 42 tests, 3 flaky, 39 stable

Multi-phase Diagnose

cargo ninety-nine diagnose finds flaky tests by stressing each test binary under multi-threaded load, then re-running only stress-failing candidates in serial isolation.

Phases

  1. Stress — for each binary that contains a selected test, run the full binary N times with --test-threads set (no --exact). Failures are parsed from libtest output and intersected with the selected set.
  2. Isolation — each candidate is run alone (--exact) N times, serially (concurrency 1, retries 0).
  3. Classify — each candidate becomes one of:
    • Contention — failed under stress, always passed alone
    • Intrinsic — failed sometimes alone
    • Broken — never passed alone
  4. Record (optional) — with --record, Intrinsic failures may be re-run under rr when available (Linux soft dependency).

Contention meaning (V1)

V1 stress is intra-binary: each test binary is exercised multi-threaded as a whole. Cross-binary workspace races are out of scope until a later runner rewrite.

Usage

# Default config ([diagnose] in .ninety-nine.toml)
cargo ninety-nine diagnose

# Override run counts
cargo ninety-nine diagnose --stress-runs 5 --isolation-runs 20

# Filter + optional rr recording
cargo ninety-nine diagnose "pkg::" --record --record-dir .ninety-nine/recordings

# JSON for CI
cargo ninety-nine diagnose -N --output json

Configuration

[diagnose]
stress_runs = 3
isolation_runs = 10
stress_threads = 0          # 0 = host parallelism
stress_timeout_secs = 300
record = false
record_dir = ".ninety-nine/recordings"
record_attempts = 10

Identity

Diagnostic rows store package + binary + test name. Bayesian scores and test_runs still use the short test name so existing filters and the scores TUI keep working.

Soft CI exit

diagnose exits 0 after a successful run even when flaky classes are found. Fail the pipeline in a wrapper if you need a hard gate.

Interactive TUI

Without -N, diagnose opens a table of CLASS | STRESS | ISOLATION | TEST | REC.

KeyAction
j / kMove
fCycle class filter (all → contention → intrinsic → broken)
EnterDetail overlay (counts + recording path)
qQuit

rr chaos mode

cargo ninety-nine diagnose --record --chaos

Requires recording enabled (--record or diagnose.record = true). Passes --chaos to rr record.

Auto-quarantine by class

[quarantine]
enabled = true
auto_quarantine = true

[quarantine.by_class]
intrinsic = true
contention = false   # default: leave load-sensitive tests visible
broken = true

Reasons are stored as auto:intrinsic, auto:broken, or auto:contention.

Multi-phase test

[detection]
multi_phase = false   # default
cargo ninety-nine test --multi-phase
cargo ninety-nine test --no-multi-phase

When enabled: diagnose stress/isolation for candidates, then Bayesian multi-run only for non-candidates.

Platform matrix (rr)

Platformrr recording
Linux + rr on PATHAvailable
Linux without rrSoft skip + install hint
macOS / WindowsSoft skip (“not supported”)

Interactive TUI

cargo ninety-nine includes a terminal user interface for browsing flakiness scores and session history. The TUI launches automatically when running status, history, or test in a terminal. Pass --non-interactive (or -N) to disable it.

Scores View

The scores view is the main screen after running test, or when running status without a test name argument. It uses bordered panels, a scrollbar, and colour-coded categories.

┌ Flaky Test Report ───────────────────────────────────────────┐
│ cargo ninety-nine | 42/42 tests shown                        │
└──────────────────────────────────────────────────────────────┘
Filter: All | Sort: P(flaky) (desc)
┌ Tests ──────────────────────────────────────────────────────┐^
│ Test                           Runs  Pass%  P(flaky) Cat.   ││
│                                                              ││
│ tests::network::retry_timeout    20  75.0%    0.250  Freq   ││
│ tests::db::concurrent_writes     20  85.0%    0.150  Mod    ││
│ tests::parser::edge_cases        20  95.0%    0.050  Occ    │█
│ tests::math::addition            20 100.0%    0.010  Stab   ││
│                                                              ││
└──────────────────────────────────────────────────────────────┘v
j/k:nav  s:sort  r:reverse  f:filter  Enter:detail  q:quit

Keybindings

KeyAction
j / Down arrowMove selection down
k / Up arrowMove selection up
sCycle sort field (Test, Runs, Pass%, P(flaky), Category)
rReverse sort order
fCycle category filter (All, Stable, Occasional, Moderate, Frequent, Critical)
EnterDrill into selected test detail
q / EscQuit
Ctrl+CQuit

Sorting

Press s to cycle through sort fields. Press r to toggle ascending/descending. The current sort field and direction are shown in the orange filter bar below the header.

Filtering

Press f to cycle through category filters. When a filter is active, only tests in that category are shown. The count in the header updates to reflect the filtered set.

Scrollbar

A vertical scrollbar appears on the right edge of the content panel. The scrollbar tracks the current selection position within the full list. The table viewport automatically follows the selected row, so scrolling through large lists works without manual page management.

Detail View

Press Enter on a test to open its detail overlay with a cyan border, showing:

  • Score summary – category, P(flaky), confidence, pass/fail rates, total runs
  • Bayesian parameters – alpha, beta, posterior mean, credible interval
  • Trend – direction (Improving/Stable/Degrading) with score delta
  • Failure patterns – correlated patterns with correlation percentage
  • Recent runs – last 10 runs with outcome, duration, and timestamp

Press Enter, q, or Esc to return to the scores list.

History View

The history view shows past detection sessions with the same bordered-panel layout:

┌ Session History ─────────────────────────────────────────────┐
│ cargo ninety-nine | 13 sessions                              │
└──────────────────────────────────────────────────────────────┘
┌ Sessions ───────────────────────────────────────────────────┐^
│ Date              Tests  Flaky  Branch           Commit     ││
│                                                              ││
│ 2026-03-23 14:39   1803      0  jason/add_tests  92d65c74   │█
│ 2026-03-23 11:34   1803      0  jason/add_tests  92d65c74   ││
│ 2026-03-18 18:39    105      0  main             5ee51231   ││
└──────────────────────────────────────────────────────────────┘v
j/k:nav  Enter:detail  q:quit

Navigate with j/k or arrow keys. Press Enter to view the test runs from that session.

Session Detail

Pressing Enter on a session opens a scrollable overlay with the same filter, sort, and reverse controls available in the scores view:

┌────────── 2026-03-26 18:38 | jason/add_tests | 47dc1602 ──────────┐
│ 9580/9580 tests | 9580 passed | 0 failed                           │
│ Filter: All | Sort: Test (asc)                                      │
│ Test                           Outcome  Duration  Retries          ^│
│                                                                    ││
│ add_slots_test                 PASS     51ms      0                ││
│ address::tests::parse_case_1  PASS     50ms      0                █│
│ address::tests::parse_case_2  PASS     50ms      0                ││
│ address::tests::roundtrip     PASS     50ms      0                ││
│                                                                    ││
│ j/k:nav  s:sort  r:reverse  f:filter  Enter/q/Esc:back            v│
└────────────────────────────────────────────────────────────────────-┘
  • Title bar – session date, branch, and commit hash
  • Summary line – filtered/total tests, passed count, failed count
  • Filter bar – current outcome filter and sort field with direction (same orange style as the scores view)
  • Test table – test name, outcome (colour-coded PASS/FAIL/TIME/PANC/SKIP), duration, retry count
  • Scrollbar – right edge of the test table with ^/v markers, tracks the selected row

Keybindings

KeyAction
j / Down arrowMove selection down
k / Up arrowMove selection up
sCycle sort field (Test, Outcome, Duration, Retries)
rReverse sort order
fCycle outcome filter (All, Pass, Fail, Timeout, Panic, Ignored)
Enter / q / EscReturn to session list
Ctrl+CQuit

Sorting

Press s to cycle through sort fields: Test name, Outcome, Duration, Retries. Press r to toggle ascending/descending. The current sort state is shown in the orange filter bar.

Filtering

Press f to cycle through outcome filters. When active, only runs with that outcome are shown. The summary line updates to show filtered/total counts.

Disabling the TUI

For CI pipelines, scripts, or piped output, use --non-interactive:

cargo ninety-nine status --non-interactive
cargo ninety-nine history --non-interactive
cargo ninety-nine test -n 10 --non-interactive

This produces the same text output as previous versions.

Visual Style

The TUI follows a panel-based layout inspired by tools like tOwl:

  • Header panel – bordered with cyan, contains the tool name and summary statistics
  • Filter bar – orange text showing current filter and sort state
  • Content panel – bordered with dark grey, contains the data table with a title
  • Scrollbar – right edge of content panel with ^/v end markers
  • Footer – keybinding hints with bold key names
  • Category colours – Stable (green), Occasional (yellow), Moderate (red), Frequent (bold red), Critical (white on red)

Terminal Requirements

The TUI uses the alternate screen buffer and raw mode via crossterm. It works in any terminal emulator that supports ANSI escape sequences. On Unix, SIGTERM and SIGHUP trigger graceful shutdown. On all platforms, the terminal state is restored on exit, including after panics.

Minimum terminal size: 60 columns by 10 rows.

Filter DSL

The filter DSL lets you select which tests to run using an expressive query language. Filters are passed as the positional argument to the test command.

cargo ninety-nine test "<filter expression>"

Syntax Overview

A filter expression is built from predicates combined with boolean operators.

Bare Words

A bare word (any identifier that is not a keyword or function call) is treated as a regex pattern matched against the test name:

# Matches any test whose name contains "network"
cargo ninety-nine test "network"

# Regex: matches test names starting with "tests::db_"
cargo ninety-nine test "tests::db_"

Function Predicates

Function predicates filter tests by metadata fields. The argument is always a single identifier.

FunctionMatches
test(pattern)Test name matches the regex pattern
package(name)Test belongs to a package whose name contains name
binary(name)Test belongs to a binary whose name contains name
kind(k)Binary kind is k – one of lib, bin, test, example
# Tests in the "my_crate" package
cargo ninety-nine test "package(my_crate)"

# Tests from test binaries only (excludes doctests, examples, etc.)
cargo ninety-nine test "kind(test)"

# Tests whose name matches a regex
cargo ninety-nine test "test(db_.*insert)"

Boolean Keywords

These keywords evaluate based on stored flakiness data:

KeywordMatches
flakyTests with at least one recorded failure, P(flaky) > 1%, and confidence at or above the threshold
quarantinedTests currently in the quarantine list
allAll tests (always true)
# Only previously-detected flaky tests
cargo ninety-nine test "flaky"

# All quarantined tests
cargo ninety-nine test "quarantined"

Note: The flaky and quarantined keywords rely on data from previous runs stored in the database. Run test at least once before using them.

Operators

Combine predicates using boolean operators:

OperatorMeaningExample
&AND – both sides must matchflaky & kind(test)
|OR – either side must matchpackage(a) | package(b)
!NOT – inverts the predicate!quarantined
( )Grouping – controls evaluation order(flaky | quarantined) & kind(test)

Operator Precedence

! (NOT) binds tightest. & and | share equal precedence and associate left to right — unlike most programming languages, & does not bind tighter than |. Use parentheses whenever an expression mixes the two.

Example: flaky | quarantined & kind(test) is parsed as (flaky | quarantined) & kind(test), because the operators apply strictly left to right. To AND first, write flaky | (quarantined & kind(test)) explicitly.

Expressions may nest ! and parentheses at most 64 levels deep; anything deeper is rejected with a parse error.

Examples

Basic Filtering

# Run all tests (no filter)
cargo ninety-nine test

# Match test names by regex
cargo ninety-nine test "integration"

# Tests in a specific package
cargo ninety-nine test "package(my_lib)"

Combining Predicates

# Flaky tests that are not quarantined
cargo ninety-nine test "flaky & !quarantined"

# Tests from either of two packages
cargo ninety-nine test "package(auth) | package(session)"

# Network tests in the integration test binary
cargo ninety-nine test "test(network) & binary(integration)"

CI-Focused Patterns

# Re-run only known flaky tests with extra iterations
cargo ninety-nine test "flaky & !quarantined" -n 50

# Run test binaries only, skip examples and benches
cargo ninety-nine test "kind(test)"

# Quarantined tests only (verify if they are still flaky)
cargo ninety-nine test "quarantined" -n 30 --confidence 0.99

# Everything except quarantined tests
cargo ninety-nine test "!quarantined"

Complex Expressions

# Flaky tests in test binaries from a specific package
cargo ninety-nine test "flaky & kind(test) & package(core)"

# Either flaky or quarantined, but only from lib targets
cargo ninety-nine test "(flaky | quarantined) & kind(lib)"

Error Handling

Invalid filter expressions produce a clear error message:

$ cargo ninety-nine test "kind(invalid)"
# Error: unknown binary kind: invalid

$ cargo ninety-nine test "unknown_func(arg)"
# Error: unknown function: unknown_func

Valid kind values are: lib, bin, test, example.

Quarantine Management

Quarantine allows you to mark flaky tests for tracking and review. Quarantined tests are stored in the SQLite database.

List Quarantined Tests

cargo ninety-nine quarantine list

Output shows test name, flakiness score, whether it was auto-quarantined, and when:

Quarantined Tests

Test                                                  Score       Auto Since
------------------------------------------------------------------------------------------------
tests::race_condition                                  33.3%        yes 2026-03-05 14:30:00
tests::timing_test                                     20.0%         no 2026-03-04 10:15:00

Add a Test to Quarantine

cargo ninety-nine quarantine add "tests::flaky_test" --reason "depends on network timing"

The --reason flag defaults to "manually quarantined" if omitted.

Remove a Test from Quarantine

cargo ninety-nine quarantine remove "tests::flaky_test"

Auto-Quarantine

Enable automatic quarantine in .ninety-nine.toml:

[quarantine]
enabled = true
auto_quarantine = true

[quarantine.threshold]
consecutive_failures = 3
failure_rate = 0.20
flakiness_score = 0.15

When auto_quarantine = true, any test detected as flaky (by the Bayesian detector) that also exceeds any of the three thresholds is automatically quarantined after a detect run.

Threshold Fields

FieldDefaultDescription
consecutive_failures3Number of consecutive failures at the tail of the run
failure_rate0.20Overall failure rate (failures / total runs)
flakiness_score0.15Bayesian P(flaky) posterior mean

A test is auto-quarantined if is_flaky(score) AND (exceeds_score OR exceeds_failures OR exceeds_rate).

JSON Output

cargo ninety-nine quarantine list --output json

Exporting Results

Export flakiness scores to various file formats for integration with CI dashboards, issue trackers, or custom tooling.

Formats

JUnit XML

cargo ninety-nine export junit results.xml

Produces standard JUnit XML. Tests with P(flaky) >= 5% are marked as failures with details in the failure message. Compatible with CI systems that parse JUnit reports (GitHub Actions, GitLab, Jenkins, etc.).

HTML Report

cargo ninety-nine export html report.html

Generates a self-contained HTML page with a styled table showing all test scores. Categories are color-coded. Suitable for sharing with teams or archiving.

CSV

cargo ninety-nine export csv results.csv

Produces a CSV file with columns:

test_name,probability_flaky,pass_rate,total_runs,consecutive_failures,category,confidence

Values containing commas, quotes, or newlines are properly escaped per RFC 4180.

JSON

cargo ninety-nine export json results.json

Produces a JSON array of flakiness score objects with full detail, including Bayesian parameters. Useful for programmatic consumption, custom dashboards, or piping into other tools.

[
  {
    "test_name": "tests::example",
    "probability_flaky": 0.167,
    "confidence": 0.95,
    "pass_rate": 0.8,
    "fail_rate": 0.2,
    "total_runs": 10,
    "consecutive_failures": 1,
    "bayesian_params": {
      "alpha": 1.0,
      "beta": 1.0,
      "posterior_mean": 0.167,
      "posterior_variance": 0.01,
      "credible_interval_lower": 0.02,
      "credible_interval_upper": 0.38
    }
  }
]

Data Source

Export uses the most recent flakiness scores stored in the SQLite database. Run test first to populate data.

CI Integration

Generating CI Workflows

Generate ready-to-use CI configuration files:

GitHub Actions

# Print to stdout
cargo ninety-nine ci generate github

# Write to file
cargo ninety-nine ci generate github .github/workflows/flaky-tests.yml

The generated workflow:

  • Runs on a weekly schedule (Monday 3:00 AM UTC) and on manual dispatch
  • Installs Rust, cargo-nextest, and cargo-ninety-nine
  • Runs flaky test detection with your configured min_runs and confidence_threshold
  • Exports JUnit XML results
  • Uploads results as a build artifact

GitLab CI

cargo ninety-nine ci generate gitlab .gitlab-ci-flaky.yml

The generated job:

  • Runs on scheduled and manual pipelines
  • Uses the rust:latest Docker image
  • Installs cargo-nextest and cargo-ninety-nine
  • Runs detection and exports JUnit XML
  • Publishes JUnit artifacts for GitLab’s test report

Failure Behaviour

Generated workflows never fail the pipeline on flaky tests: the GitHub Actions job carries continue-on-error: true and the GitLab job carries allow_failure: true, so detection results arrive as reports rather than as red builds. Remove those lines from the generated file if you would rather have flaky detections break the build.

Manual CI Setup

If you prefer to configure CI manually:

# Install
cargo install cargo-ninety-nine

# Run detection
cargo ninety-nine test -n 20 --confidence 0.95

# Export for CI test report parsing
cargo ninety-nine export junit flaky-results.xml

Environment Detection

cargo ninety-nine automatically detects the CI environment:

Environment VariableDetected Provider
GITHUB_ACTIONSGitHub Actions
GITLAB_CIGitLab CI
JENKINS_URLJenkins
CIRCLECICircleCI
BUILDKITEBuildkite
TF_BUILDAzure DevOps

The detected provider is stored with each test run, enabling environmental pattern analysis (CI vs local failure rate differences).

Best Practices

Guidance for getting the most out of cargo-ninety-nine in real-world workflows.

Iteration Strategy

Choosing min_runs

The number of iterations directly impacts detection accuracy:

IterationsUse CaseConfidence
5–10Quick smoke checkLow — high false positive rate
10–20Daily developmentModerate — good for known-flaky tests
20–50Pre-merge gateHigh — reliable classification
50–100Baseline establishmentVery high — suitable for initial assessment

Note: The Bayesian detector uses a uniform Beta(1,1) prior. With fewer than 10 runs, the prior dominates the posterior and scores are unreliable. The default min_runs = 10 is a minimum for meaningful results.

Building a Baseline

When first adopting cargo-ninety-nine, establish a baseline:

# Run with higher iterations to build confidence
cargo ninety-nine test -n 50

# Review the full status
cargo ninety-nine status

Subsequent runs build on the historical data, so daily runs with fewer iterations (10–20) are sufficient once the baseline exists.

Quarantine Strategy

When to Quarantine

Quarantine a test when:

  • It blocks CI pipelines with intermittent failures
  • It has a Moderate or higher flakiness category (>= 0.05)
  • It has 3+ consecutive failures with no code changes

Do not quarantine a test when:

  • It consistently fails — that is a real bug, not flakiness
  • It just started failing after a recent change — investigate the change first
  • The flakiness score is Occasional (<0.05) — monitor it instead

Auto-Quarantine Thresholds

If enabling auto-quarantine, tune the thresholds to avoid over-quarantining:

[quarantine]
auto_quarantine = true

[quarantine.threshold]
consecutive_failures = 5    # stricter than default 3
failure_rate = 0.30         # stricter than default 0.20
flakiness_score = 0.20      # stricter than default 0.15

Start strict and relax gradually based on your team’s tolerance.

Quarantine Review Cadence

Quarantine entries persist until removed, so agree a review cadence with your team — fortnightly works well — and walk the list each time:

  1. cargo ninety-nine quarantine list — see all quarantined tests
  2. Re-run quarantined tests: cargo ninety-nine test "quarantined()"
  3. Remove fixed tests: cargo ninety-nine quarantine remove <test_name>

CI Integration

Run cargo-ninety-nine as a separate CI job that does not block merges:

# GitHub Actions example
flaky-detection:
  runs-on: ubuntu-latest
  if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
  steps:
    - uses: actions/checkout@v4
    - uses: dtolnay/rust-toolchain@stable
    - run: cargo install cargo-ninety-nine
    - run: cargo ninety-nine test -n 20
    - run: cargo ninety-nine export junit report.xml
    - uses: actions/upload-artifact@v4
      with:
        name: flaky-report
        path: report.xml

Warning: Avoid running flaky detection on every push. It adds significant CI time and produces noise. Schedule it nightly or weekly.

Shared Storage in CI

For teams wanting to accumulate results across CI runs, use PostgreSQL:

[storage]
backend = "Postgres"

[storage.postgres]
connection_string = "host=db.internal dbname=ninety_nine user=ci password=${NN_PG_PASSWORD}"
pool_size = 4

Pass the password via environment variable in CI secrets.

Performance Tuning

Parallel Execution

The parallel_runs setting controls how many tests execute concurrently:

[detection]
parallel_runs = 4  # default: 3

Guidelines:

  • Local development: set to number of CPU cores / 2
  • CI (shared runners): set to 1–2 to avoid resource contention
  • Dedicated CI machines: set to CPU count

Reducing Test Discovery Time

Test discovery runs cargo test --no-run, which may trigger a full build. To speed this up:

  • Keep your build cache warm (target/ directory)
  • Use incremental compilation
  • Consider running detection on a subset: cargo ninety-nine test "package(critical_module)"

Database Performance

SQLite (default):

  • Excellent for single-machine use
  • WAL mode is enabled automatically for concurrent reads
  • Keep the database on a local SSD, not a network filesystem

PostgreSQL:

  • Better for shared/team use and large datasets
  • Tune pool_size based on concurrent CI jobs
  • Use connection pooling (PgBouncer) for high-concurrency environments

Interpreting Results

Understanding Bayesian Scores

ScoreMeaningAction
< 0.01Stable — no flakiness detectedNo action needed
0.01 – 0.05Occasional — rare failuresMonitor, investigate if trending up
0.05 – 0.15Moderate — noticeable flakinessInvestigate root cause, consider quarantine
0.15 – 0.30Frequent — regular failuresFix or quarantine immediately
>= 0.30Critical — failing more than passingUrgent fix required

Acting on Patterns

PatternCommon CausesRemediation
Time-of-dayScheduled jobs competing for resources, time-sensitive assertionsRemove time dependencies, use faketime in tests
EnvironmentalMissing dependencies in CI, different OS behaviorEnsure CI mirrors local dev environment, use containers
RandomRace conditions, shared mutable state, non-deterministic orderingAdd synchronization, use deterministic seeds, isolate test state

Confidence and Sample Size

Low confidence means the credible interval is wide — the true flakiness probability could be much higher or lower than the point estimate. Before acting on a score:

  • High confidence (>0.95): Act on the score directly
  • Moderate confidence (0.80–0.95): Run more iterations to confirm
  • Low confidence (<0.80): Too few runs to draw conclusions, need more data

Migrating from SQLite to PostgreSQL

  1. Export current data:

    cargo ninety-nine export json current-data.json
    
  2. Update configuration:

    [storage]
    backend = "Postgres"
    
    [storage.postgres]
    connection_string = "host=localhost dbname=ninety_nine"
    pool_size = 4
    
  3. Run a fresh detection pass to populate the new database:

    cargo ninety-nine test -n 20
    

Note: There is no automatic migration tool. Historical run data is not transferred — only flakiness scores and quarantine state can be preserved via JSON export. The new database will build fresh statistics from subsequent runs.

Team Workflows

Flaky Test Triage

Establish a regular triage process:

  1. Weekly: Review cargo ninety-nine status output
  2. Per-sprint: Assign Moderate+ flaky tests to developers
  3. Per-release: Clear all Frequent/Critical tests before release

Filter Expressions for Triage

# Show only flaky tests
cargo ninety-nine test "flaky()"

# Focus on a specific package
cargo ninety-nine test "package(api) & flaky()"

# Exclude already-quarantined tests
cargo ninety-nine test "flaky() & !quarantined()"

Using JSON Output for Automation

# Export status as JSON for dashboards
cargo ninety-nine --output json status > flaky-status.json

# Parse with jq
cat flaky-status.json | jq '.[] | select(.probability_flaky > 0.1)'

Troubleshooting

Common issues and their solutions when using cargo-ninety-nine.

Installation Issues

no test runner available

error: no test runner available: install cargo-nextest or use cargo test

Cause: Neither cargo-nextest nor cargo test was found on the system PATH.

Solutions:

  1. Install cargo-nextest: cargo install cargo-nextest
  2. Verify Rust toolchain is installed: rustup show
  3. Ensure ~/.cargo/bin is in your PATH

binary discovery failed

error: binary discovery failed: ...

Cause: cargo test --no-run failed to compile or list test binaries.

Solutions:

  1. Run cargo test --no-run manually to see the full compiler output
  2. Fix any compilation errors in your project
  3. Ensure you are running from the project root (or use --project-dir)

Configuration Issues

failed to parse config

error: failed to parse config: ...

Cause: The .ninety-nine.toml file contains invalid TOML syntax or unrecognized fields.

Solutions:

  1. Validate your TOML syntax: check for unclosed quotes, missing brackets, or incorrect indentation
  2. Re-generate a fresh config: cargo ninety-nine init --force
  3. Compare against the default config shown in the Configuration Reference

postgres backend selected but no config

error: invalid configuration: postgres backend selected but no [storage.postgres] config provided

Cause: storage.backend is set to "Postgres" but the [storage.postgres] section is missing.

Solution: Add the PostgreSQL configuration:

[storage]
backend = "Postgres"

[storage.postgres]
connection_string = "host=localhost dbname=ninety_nine user=postgres"
pool_size = 4

Test Execution Issues

Tests not discovered

Symptom: cargo ninety-nine test reports 0 tests found.

Causes and solutions:

  1. No test targets: Ensure your project has #[test] functions or files in tests/
  2. Filter too restrictive: Remove or widen your filter expression
  3. Wrong project directory: Use --project-dir /path/to/project
  4. Benchmark-only binaries: Only #[test] functions are discovered, not benchmarks

Test timeouts

Symptom: Tests that normally pass are reported as Timeout.

Causes and solutions:

  1. Default timeout too low: The default is 300 seconds. For long-running tests, increase it in config:
    [detection]
    # Timeout is controlled via the execution config
    
  2. CI resource constraints: CI environments often have fewer resources. Check memory_gb and cpu_count in the environment report
  3. Test contention: Reduce parallel_runs to lower resource contention:
    [detection]
    parallel_runs = 1
    

All tests show as flaky

Symptom: Every test gets a non-trivial flakiness score.

Causes and solutions:

  1. Too few iterations: With only a few runs, the Bayesian prior has outsized influence. Increase min_runs:
    [detection]
    min_runs = 20
    
  2. Confidence threshold too low: Raise the threshold to require stronger evidence:
    [detection]
    confidence_threshold = 0.99
    
  3. Systemic failures: If tests are failing due to environment issues (missing database, network), fix the root cause rather than tuning thresholds

Storage Issues

SQLite database locked

Symptom: storage error: database is locked

Causes and solutions:

  1. Concurrent access: Another cargo-ninety-nine process may be running. WAL mode should handle most concurrent access, but heavy parallel writes can still lock
  2. NFS/network filesystem: SQLite does not work reliably over network filesystems. Use a local path or switch to PostgreSQL
  3. Stale lock: If the process crashed, the lock file may remain. Delete ninety-nine.db-wal and ninety-nine.db-shm next to the database

PostgreSQL connection failures

Symptom: postgres storage error: connection refused or similar

Solutions:

  1. Verify PostgreSQL is running: pg_isready
  2. Check connection string format: host=localhost port=5432 dbname=ninety_nine user=postgres password=...
  3. Ensure the database exists: createdb ninety_nine
  4. Check network/firewall settings for remote connections
  5. Verify pool size is reasonable for your connection limits

Data retention and database size

Symptom: Database growing too large.

Solution: Configure retention_days to automatically purge old data:

[storage]
retention_days = 30  # default: 90

Data is purged at the end of each test run.

Filter DSL Issues

filter parse error

Symptom: error: filter parse error: ...

Common mistakes:

  1. Missing parentheses on predicates: Use flaky() not flaky
  2. Wrong operator syntax: Use & for AND, | for OR, ! for NOT
  3. Unclosed parentheses: Ensure every ( has a matching )
  4. Invalid regex in test(): The pattern must be a valid Rust regex

Valid examples:

flaky()
test(.*timeout.*)
package(auth) & !quarantined()
(flaky() | test(.*race.*)) & package(core)

CI Integration Issues

Workflow not triggering

Symptom: Generated CI workflow never runs.

Solutions:

  • GitHub Actions: Ensure the workflow file is at .github/workflows/. Check that scheduled triggers are on the default branch
  • GitLab CI: Ensure the pipeline is configured to run on schedules. Check that rules allow scheduled execution

CI environment not detected

Symptom: is_ci shows false in CI.

Cause: The CI provider’s environment variable is not set or not recognized.

Recognized variables:

VariableProvider
GITHUB_ACTIONSGitHub Actions
GITLAB_CIGitLab CI
JENKINS_URLJenkins
CIRCLECICircleCI
TF_BUILDAzure DevOps
BUILDKITEBuildkite

Export Issues

Empty export files

Symptom: Export produces a file with no test data.

Cause: No flakiness scores have been computed yet.

Solution: Run cargo ninety-nine test at least once before exporting. Scores are computed and stored during the test run.

JUnit XML not recognized

Symptom: CI system does not parse the JUnit XML output.

Solution: Ensure the export path matches what your CI expects. For GitHub Actions:

- uses: dorny/test-reporter@v1
  with:
    artifact: ninety-nine-report
    name: Flaky Tests
    path: report.xml
    reporter: java-junit

Getting More Information

Enable verbose output for detailed tracing:

cargo ninety-nine --verbose test

This sets the tracing subscriber to debug level, showing:

  • Binary discovery details
  • Test listing parsing
  • Individual test execution results
  • Storage operations
  • Bayesian computation details

Types Reference

This chapter documents the core types used throughout cargo-ninety-nine.

Test Execution

TestRun

A single test execution record with full metadata.

#![allow(unused)]
fn main() {
pub struct TestRun {
    pub id: Uuid,
    pub test_name: TestName,
    pub test_path: PathBuf,
    pub outcome: TestOutcome,
    pub duration: Duration,
    pub timestamp: DateTime<Utc>,
    pub commit_hash: String,
    pub branch: String,
    pub environment: TestEnvironment,
    pub retry_count: u32,
    pub error_message: Option<String>,
    pub stack_trace: Option<String>,
}
}
FieldTypeDescription
idUuidUnique identifier for this run
test_nameTestNameType-safe test name
test_pathPathBufPath to the test binary
outcomeTestOutcomeResult of the execution
durationDurationWall-clock execution time
timestampDateTime<Utc>When the run occurred
commit_hashStringGit commit at time of run
branchStringGit branch at time of run
environmentTestEnvironmentExecution environment details
retry_countu32Number of retries before this result
error_messageOption<String>Failure message, if any
stack_traceOption<String>Stack trace on panic/failure

TestOutcome

Classification of a test execution result.

#![allow(unused)]
fn main() {
pub enum TestOutcome {
    Passed,
    Failed,
    Ignored,
    Timeout,
    Panic,
}
}
VariantDescription
PassedTest completed successfully (exit code 0)
FailedTest assertion failed (non-zero exit, no panic)
IgnoredTest marked with #[ignore]
TimeoutTest exceeded the configured timeout
PanicTest panicked (detected via panicked at in output)

Implements Display and FromStr for serialization. Display values: "passed", "failed", "ignored", "timeout", "panic".

TestEnvironment

Captures the environment where tests execute, used for pattern correlation.

#![allow(unused)]
fn main() {
pub struct TestEnvironment {
    pub os: String,
    pub rust_version: String,
    pub cpu_count: u32,
    pub memory_gb: f64,
    pub is_ci: bool,
    pub ci_provider: Option<String>,
}
}

Auto-detected at runtime from the host system. The is_ci field is inferred from environment variables (GITHUB_ACTIONS, GITLAB_CI, JENKINS_URL, CIRCLECI, TF_BUILD, BUILDKITE).

TestName

Newtype wrapper providing type-safe test names. Prevents accidental confusion with branch names, commit hashes, or other string fields.

#![allow(unused)]
fn main() {
pub struct TestName(String);
}

Conversions:

  • From<String>, From<&str> — construct from strings
  • Deref<Target = str> — borrow as &str
  • AsRef<str> — reference conversion
  • Display — format for output
  • PartialEq<str>, PartialEq<&str> — compare with string values

Methods:

  • into_inner(self) -> String — consume and return the inner string

Flakiness Detection

FlakinessScore

The primary output of the Bayesian detection engine. Contains the computed probability that a test is flaky along with statistical parameters.

#![allow(unused)]
fn main() {
pub struct FlakinessScore {
    pub test_name: TestName,
    pub probability_flaky: f64,
    pub confidence: f64,
    pub pass_rate: f64,
    pub fail_rate: f64,
    pub total_runs: u64,
    pub consecutive_failures: u32,
    pub last_updated: DateTime<Utc>,
    pub bayesian_params: BayesianParams,
}
}
FieldTypeRangeDescription
probability_flakyf64[0.0, 1.0]Posterior mean P(failure)
confidencef64[0.0, 1.0]1 - credible interval width (higher = more certain)
pass_ratef64[0.0, 1.0]Fraction of runs that passed
fail_ratef64[0.0, 1.0]Fraction of runs that failed (= 1 - pass_rate)
total_runsu64—Number of non-ignored executions
consecutive_failuresu32—Trailing failure streak count

BayesianParams

Full Bayesian computation state stored for auditability.

#![allow(unused)]
fn main() {
pub struct BayesianParams {
    pub alpha: f64,
    pub beta: f64,
    pub posterior_mean: f64,
    pub posterior_variance: f64,
    pub credible_interval_lower: f64,
    pub credible_interval_upper: f64,
}
}
FieldDescription
alphaBeta distribution shape parameter (prior + failures)
betaBeta distribution shape parameter (prior + passes)
posterior_meanalpha / (alpha + beta)
posterior_variance(alpha * beta) / ((alpha + beta)^2 * (alpha + beta + 1))
credible_interval_lower2.5th percentile of posterior Beta distribution
credible_interval_upper97.5th percentile of posterior Beta distribution

FlakinessCategory

Human-readable classification of flakiness severity.

#![allow(unused)]
fn main() {
pub enum FlakinessCategory {
    Stable,
    Occasional,
    Moderate,
    Frequent,
    Critical,
}
}
CategoryScore RangeConsole Color
Stable< 0.01Green
Occasional0.01 – 0.05Yellow
Moderate0.05 – 0.15Orange
Frequent0.15 – 0.30Red
Critical>= 0.30Dark Red

Methods:

  • from_score(score: f64) -> Self — classify a probability value
  • label(&self) -> &'static str — human-readable label

Sessions

ActiveSession

Represents a running test session. Created via start(); to_run_session() produces the storable RunSession snapshot while the session stays live.

#![allow(unused)]
fn main() {
pub struct ActiveSession { /* private fields */ }
}

Methods:

MethodSignatureDescription
startfn start(commit_hash: &str, branch: &str) -> SelfCreate a new session
idfn id(&self) -> &UuidGet session UUID
to_run_sessionfn to_run_session(&self) -> RunSessionConvert to storable form (borrowing)

RunSession

A session record suitable for storage. May represent a running or completed session.

#![allow(unused)]
fn main() {
pub struct RunSession {
    pub id: Uuid,
    pub started_at: DateTime<Utc>,
    pub finished_at: Option<DateTime<Utc>>,
    pub test_count: u32,
    pub flaky_count: u32,
    pub commit_hash: String,
    pub branch: String,
}
}

QuarantineEntry

A quarantined test record.

#![allow(unused)]
fn main() {
pub struct QuarantineEntry {
    pub test_name: TestName,
    pub quarantined_at: DateTime<Utc>,
    pub reason: String,
    pub flakiness_score: f64,
    pub auto_quarantined: bool,
}
}

Analysis

TrendDirection

Direction of flakiness change over time.

#![allow(unused)]
fn main() {
pub enum TrendDirection {
    Improving,
    Stable,
    Degrading,
}
}

A delta exceeding 0.05 (5%) triggers Improving (decreased flakiness) or Degrading (increased flakiness). Otherwise Stable.

TrendSummary

Trend analysis result comparing recent vs. historical flakiness.

#![allow(unused)]
fn main() {
pub struct TrendSummary {
    pub test_name: TestName,
    pub direction: TrendDirection,
    pub recent_score: f64,
    pub previous_score: f64,
    pub score_delta: f64,
    pub window_runs: u64,
}
}

FailurePattern

A detected failure pattern with correlation strength.

#![allow(unused)]
fn main() {
pub struct FailurePattern {
    pub pattern_type: PatternType,
    pub occurrences: u32,
    pub correlation: f64,
    pub examples: Vec<String>,
}
}

PatternType

Classification of detected failure patterns.

#![allow(unused)]
fn main() {
pub enum PatternType {
    TimeOfDay,
    Environmental,
    Random,
}
}
VariantTriggerDescription
TimeOfDayFailure concentration > 3x expected in a specific hourFailures cluster at a particular time
EnvironmentalCI vs. local failure rate difference > 15%Environment-specific failures
RandomNo pattern detectedFailures appear randomly distributed

Bayesian Detector API

The BayesianDetector is the core statistical engine that computes flakiness probabilities using Beta-Binomial conjugate inference.

BayesianDetector

#![allow(unused)]
fn main() {
pub struct BayesianDetector {
    prior_alpha: f64,     // default: 1.0 (uniform prior)
    prior_beta: f64,      // default: 1.0 (uniform prior)
    confidence_threshold: f64,
}
}

Constructor

#![allow(unused)]
fn main() {
pub const fn new(confidence_threshold: f64) -> Self
}

Creates a detector with a uniform Beta(1, 1) prior and the given confidence threshold.

ParameterTypeDescription
confidence_thresholdf64Minimum confidence required to classify a test as flaky (typically 0.95)

Methods

calculate_flakiness_score

#![allow(unused)]
fn main() {
pub fn calculate_flakiness_score(
    &self,
    test_name: &str,
    runs: &[TestRun],
) -> FlakinessScore
}

Computes a FlakinessScore from a set of test runs using Bayesian inference.

Algorithm:

  1. Count passes and failures from runs (ignoring Ignored outcomes; Failed, Panic, and Timeout all count as failures)
  2. Update the Beta prior: alpha = prior_alpha + failures, beta = prior_beta + passes
  3. Compute posterior mean: alpha / (alpha + beta)
  4. Compute posterior variance: (alpha * beta) / ((alpha + beta)^2 * (alpha + beta + 1))
  5. Compute 95% credible interval via Beta::inverse_cdf(0.025) and Beta::inverse_cdf(0.975)
  6. Confidence = 1.0 - (upper - lower) (narrower interval = higher confidence)
  7. Count consecutive trailing failures

Returns: A FlakinessScore with all computed fields populated.

is_flaky

#![allow(unused)]
fn main() {
pub fn is_flaky(&self, score: &FlakinessScore) -> bool
}

Determines whether a test should be classified as flaky.

Criteria: Returns true when both conditions are met:

  • score.probability_flaky > 0.01 — non-trivial failure probability
  • score.confidence >= confidence_threshold — sufficient statistical certainty

Internal Functions

These are not public but documented for understanding the algorithm:

FunctionPurpose
count_outcomes(runs)Counts passes and failures, ignoring Ignored outcomes
count_consecutive_trailing_failures(runs)Counts failures from the end of the run list until a pass is encountered
credible_interval(alpha, beta)Computes 2.5th–97.5th percentile interval using the statrs Beta distribution

Usage Example

#![allow(unused)]
fn main() {
use cargo_ninety_nine::detector::BayesianDetector;

let detector = BayesianDetector::new(0.95);
let score = detector.calculate_flakiness_score("tests::my_test", &runs);

if detector.is_flaky(&score) {
    println!("{} is flaky (P={:.2})", score.test_name, score.probability_flaky);
}
}

Statistical Properties

The detector guarantees the following properties (verified by property-based tests):

  • probability_flaky is always in [0.0, 1.0]
  • confidence is always in [0.0, 1.0]
  • More failures relative to passes always produces a higher probability_flaky
  • The credible interval is always valid: lower >= 0.0, upper <= 1.0, lower <= upper

See the Bayesian Detection reference for the full mathematical model.

Analysis API

The analysis module provides three capabilities: failure rate computation, trend detection, and failure pattern recognition.

Failure Rate

failure_rate

#![allow(unused)]
fn main() {
pub fn failure_rate(runs: &[&TestRun]) -> f64
}

Computes the fraction of runs that resulted in a failure outcome.

OutcomeCounted as
PassedNon-failure
IgnoredNon-failure
FailedFailure
PanicFailure
TimeoutFailure

Returns: A value in [0.0, 1.0]. Returns 0.0 for empty input.

Trend Detection

calculate_trend

#![allow(unused)]
fn main() {
pub fn calculate_trend(
    test_name: &str,
    runs: &[TestRun],
    window: u32,
) -> Option<TrendSummary>
}

Analyzes the direction of flakiness change by comparing recent runs against historical runs.

ParameterTypeDescription
test_name&strName of the test being analyzed
runs&[TestRun]Runs to analyze (newest first not required)
windowu32Maximum number of runs to consider

Algorithm:

  1. Takes up to window most recent runs
  2. Requires a minimum of 4 runs; returns None otherwise
  3. Splits runs at the midpoint into “recent” (first half) and “previous” (second half)
  4. Computes failure rate for each half
  5. Classifies direction based on the delta:
    • Delta < -0.05 (5% improvement): Improving
    • Delta > +0.05 (5% regression): Degrading
    • Otherwise: Stable

Returns: Some(TrendSummary) with direction, scores, and delta, or None if insufficient data.

Duration Regression

detect_duration_regressions

#![allow(unused)]
fn main() {
pub fn detect_duration_regressions(
    test_name: &str,
    runs: &[TestRun],
    min_history: usize,
    threshold: RegressionThreshold,
) -> Option<DurationRegression>
}

Detects if the most recent test run is significantly slower than historical runs.

ParameterTypeDescription
test_name&strName of the test
runs&[TestRun]Runs ordered newest first (at least min_history required)
min_historyusizeMinimum number of runs required
thresholdRegressionThresholdHow the latest duration is judged against the baseline
#![allow(unused)]
fn main() {
pub enum RegressionThreshold {
    StdDevs(f64),
    Multiplier(f64),
}
}

Algorithm:

  1. Returns None if fewer than min_history runs
  2. Computes mean and standard deviation of the historical durations (the latest run is excluded from the baseline)
  3. Returns None if the effective standard deviation is near zero
  4. StdDevs(z) flags when the latest duration exceeds mean + z × std_dev; Multiplier(m) flags when it exceeds mean × m
  5. Returns DurationRegression with deviation statistics

DurationRegression

#![allow(unused)]
fn main() {
pub struct DurationRegression {
    pub test_name: TestName,
    pub current_ms: f64,
    pub mean_ms: f64,
    pub std_dev_ms: f64,
    pub deviation_factor: f64,
}
}
FieldDescription
current_msDuration of the latest run in milliseconds
mean_msHistorical mean duration in milliseconds
std_dev_msHistorical standard deviation in milliseconds
deviation_factorHow many standard deviations above mean: (current - mean) / std_dev

Pattern Detection

detect_patterns

#![allow(unused)]
fn main() {
pub fn detect_patterns(runs: &[TestRun]) -> Vec<FailurePattern>
}

Scans test runs for recurring failure patterns. Returns all detected patterns.

Detection strategies:

Time-of-Day Pattern

Bins all failures by hour (0–23). If any hour contains more than 3x the expected random concentration, a TimeOfDay pattern is reported.

  • Correlation value: ratio of peak-hour failures to total failures
  • Examples: formatted as "Hour HH: N failures" for the peak hour

Environmental Pattern

Compares failure rates between CI and local environments. If the difference exceeds 15%, an Environmental pattern is reported.

  • Correlation value: absolute difference between CI and local failure rates
  • Examples: formatted as "CI failure rate: X%, local: Y%"

Random Fallback

If no time-of-day or environmental pattern is detected, a Random pattern is returned with correlation 0.0, indicating failures appear uniformly distributed.

Usage Example

#![allow(unused)]
fn main() {
use cargo_ninety_nine::analysis::{calculate_trend, detect_patterns};
use cargo_ninety_nine::analysis::duration::detect_duration_regressions;

// Trend analysis
if let Some(trend) = calculate_trend("tests::my_test", &runs, 100) {
    println!("Trend: {} (delta: {:.2})", trend.direction, trend.score_delta);
}

// Duration regression
use cargo_ninety_nine::analysis::duration::RegressionThreshold;
if let Some(reg) =
    detect_duration_regressions("tests::my_test", &runs, 5, RegressionThreshold::StdDevs(2.0))
{
    println!("SLOW: {}ms vs mean {}ms ({:.1}x std_dev)",
        reg.current_ms, reg.mean_ms, reg.deviation_factor);
}

// Pattern detection
for pattern in detect_patterns(&runs) {
    println!("Pattern: {} (correlation: {:.2})", pattern.pattern_type, pattern.correlation);
}
}

Filter DSL API

The filter module implements a small domain-specific language for selecting tests. It consists of four stages: tokenization, parsing, context building, and evaluation.

Pipeline Overview

input string → tokenize() → parse() → FilterExpr (AST)
                                            ↓
test metadata + EvalContext → eval() → bool (match/no-match)

Compilation

compile_filter

#![allow(unused)]
fn main() {
pub fn compile_filter(input: &str) -> Result<FilterExpr, NinetyNineError>
}

Compiles a filter expression string into an AST. This is the primary entry point for the filter DSL.

Errors: Returns NinetyNineError::FilterParse if the input is syntactically invalid.

build_eval_context

#![allow(unused)]
fn main() {
pub async fn build_eval_context(
    storage: &impl Storage,
    confidence: f64,
) -> Result<EvalContext, NinetyNineError>
}

Pre-loads data from storage needed to evaluate predicates like flaky() and quarantined().

The context contains:

  • Flaky test set: tests where probability_flaky > 0.01 and confidence >= threshold
  • Quarantined test set: all currently quarantined tests

AST Types

FilterExpr

The abstract syntax tree for filter expressions.

#![allow(unused)]
fn main() {
pub enum FilterExpr {
    And(Vec<FilterExpr>),
    Or(Vec<FilterExpr>),
    Not(Box<FilterExpr>),
    Predicate(Predicate),
}
}

Predicate

Leaf-level predicates that match against test metadata.

#![allow(unused)]
fn main() {
pub enum Predicate {
    Test(Regex),       // match test name against regex
    Package(String),   // match package name
    Binary(String),    // match binary name
    Kind(BinaryKind),  // match binary kind (lib, bin, test, example)
    Flaky,             // test is currently flaky
    Quarantined,       // test is currently quarantined
    All,               // matches everything
}
}

Tokenizer

tokenize

#![allow(unused)]
fn main() {
pub fn tokenize(input: &str) -> Result<Vec<Token>, NinetyNineError>
}

Splits input into tokens.

Token

#![allow(unused)]
fn main() {
pub enum Token {
    Ident(String),
    LParen,
    RParen,
    And,
    Or,
    Not,
    Equals,
}
}

Parser

parse

#![allow(unused)]
fn main() {
pub fn parse(tokens: Vec<Token>) -> Result<FilterExpr, NinetyNineError>
}

Recursive descent parser with this precedence (lowest to highest):

  1. Or (|) — binary, left-associative
  2. And (&) — binary, left-associative
  3. Not (!) — unary prefix
  4. Primary — parenthesized expressions or predicate calls

Predicate syntax:

test(pattern)       — regex match on test name
package(name)       — exact match on package name
binary(name)        — exact match on binary name
kind(lib|bin|test|example) — match binary kind
flaky()             — test is flaky
quarantined()       — test is quarantined
all()               — match everything

Evaluator

eval

#![allow(unused)]
fn main() {
pub fn eval(
    expr: &FilterExpr,
    meta: &TestMetadata,
    ctx: &EvalContext,
) -> bool
}

Evaluates a compiled FilterExpr against a test’s metadata using a pre-built evaluation context.

TestMetadata

#![allow(unused)]
fn main() {
pub struct TestMetadata {
    pub name: String,
    pub package_name: String,
    pub binary_name: String,
    pub kind: BinaryKind,
}
}

Constructed from a TestCase at evaluation time.

EvalContext

#![allow(unused)]
fn main() {
pub struct EvalContext {
    pub flaky_tests: HashSet<String>,
    pub quarantined_tests: HashSet<String>,
}
}

Pre-computed sets for O(1) predicate evaluation.

Examples

# All flaky tests in the "auth" package
package(auth) & flaky()

# Everything except quarantined tests
!quarantined()

# Tests matching a pattern OR known flaky tests
test(.*timeout.*) | flaky()

# Complex composition
(package(api) | package(core)) & !quarantined() & flaky()

See the Filter DSL Guide for usage examples from the CLI.

Storage API

The storage module provides an async trait abstraction over test result persistence, with SQLite and PostgreSQL backends.

Storage Trait

#![allow(unused)]
fn main() {
pub trait Storage: Send + Sync {
    async fn store_session(&self, session: &RunSession) -> Result<(), NinetyNineError>;
    async fn finish_session(&self, session_id: &Uuid, test_count: u32, flaky_count: u32) -> Result<(), NinetyNineError>;
    async fn store_test_run(&self, run: &TestRun, session_id: &Uuid) -> Result<(), NinetyNineError>;
    async fn store_flakiness_score(&self, score: &FlakinessScore) -> Result<(), NinetyNineError>;
    async fn get_test_runs(&self, test_name: &str, limit: u32) -> Result<Vec<TestRun>, NinetyNineError>;
    async fn get_recent_sessions(&self, limit: u32) -> Result<Vec<RunSession>, NinetyNineError>;
    async fn get_all_scores(&self) -> Result<Vec<FlakinessScore>, NinetyNineError>;
    async fn get_score(&self, test_name: &str) -> Result<Option<FlakinessScore>, NinetyNineError>;
    async fn quarantine_test(&self, test_name: &str, reason: &str, score: f64, auto: bool) -> Result<(), NinetyNineError>;
    async fn unquarantine_test(&self, test_name: &str) -> Result<(), NinetyNineError>;
    async fn get_quarantined_tests(&self) -> Result<Vec<QuarantineEntry>, NinetyNineError>;
    async fn is_quarantined(&self, test_name: &str) -> Result<bool, NinetyNineError>;
    async fn get_session_runs(&self, session_id: &Uuid) -> Result<Vec<TestRun>, NinetyNineError>;
    async fn purge_older_than(&self, days: u32) -> Result<u64, NinetyNineError>;
}
}

All methods are async to support both synchronous (SQLite via spawn_blocking) and natively async (PostgreSQL) backends.

Method Reference

MethodDescription
store_sessionPersists a new run session
finish_sessionMarks a session as complete with final test/flaky counts
store_test_runStores a single test execution result
store_flakiness_scoreUpserts a computed flakiness score (INSERT OR REPLACE)
get_test_runsRetrieves runs for a test, most recent first
get_recent_sessionsLists recent sessions, most recent first
get_all_scoresReturns all scores ordered by probability_flaky descending
get_scoreReturns the score for a specific test, or None
quarantine_testAdds a test to quarantine with reason and score
unquarantine_testRemoves a test from quarantine
get_quarantined_testsLists all quarantined tests
is_quarantinedChecks if a specific test is quarantined
get_session_runsRetrieves all test runs for a given session, ordered by test name
purge_older_thanDeletes test runs older than N days, returns count deleted

StorageBackend Enum

#![allow(unused)]
fn main() {
pub enum StorageBackend {
    Sqlite(SqliteStorage),
    Postgres(PostgresStorage),
}
}

Delegates all Storage trait methods to the underlying backend.

Factory Function

open_storage

#![allow(unused)]
fn main() {
pub async fn open_storage(config: &Config) -> Result<StorageBackend, NinetyNineError>
}

Opens the configured storage backend:

  • StorageBackendType::Sqlite — Opens (or creates) a SQLite database at the configured path, defaulting to $XDG_DATA_HOME/ninety-nine/ninety-nine.db
  • StorageBackendType::Postgres — Connects to PostgreSQL using the configured connection string and pool size

Errors: Returns InvalidConfig if Postgres is selected but no [storage.postgres] configuration is provided.

SQLite Backend

SqliteStorage

#![allow(unused)]
fn main() {
pub struct SqliteStorage { /* private fields */ }
}

Constructor:

#![allow(unused)]
fn main() {
pub fn open(db_path: &Path) -> Result<Self, NinetyNineError>
}

Opens the database, creates parent directories if needed, enables WAL mode for concurrent reads/writes, and runs schema migrations.

Features:

  • WAL (Write-Ahead Logging) mode for better concurrent access
  • Automatic schema creation on first open
  • Bundled SQLite via rusqlite (no system dependency required)

Default location: $XDG_DATA_HOME/ninety-nine/ninety-nine.db

  • Linux: ~/.local/share/ninety-nine/ninety-nine.db
  • macOS: ~/Library/Application Support/ninety-nine/ninety-nine.db

PostgreSQL Backend

PostgresStorage

#![allow(unused)]
fn main() {
pub struct PostgresStorage { /* private fields */ }
}

Constructor:

#![allow(unused)]
fn main() {
pub async fn connect(
    connection_string: &str,
    pool_size: u32,
) -> Result<Self, NinetyNineError>
}

Connects to PostgreSQL and initializes the schema. Uses deadpool-postgres for connection pooling.

ParameterTypeDescription
connection_string&strPostgreSQL connection URL
pool_sizeu32Maximum number of pooled connections

Features:

  • Connection pooling via deadpool-postgres
  • Automatic schema creation on first connect
  • Natively async operations (no spawn_blocking needed)

Database Schema

Both backends share the same logical schema with four tables:

-- Test execution sessions
CREATE TABLE sessions (
    id TEXT PRIMARY KEY,
    started_at TEXT NOT NULL,
    finished_at TEXT,
    test_count INTEGER NOT NULL DEFAULT 0,
    flaky_count INTEGER NOT NULL DEFAULT 0,
    commit_hash TEXT NOT NULL,
    branch TEXT NOT NULL
);

-- Individual test run results
CREATE TABLE test_runs (
    id TEXT PRIMARY KEY,
    session_id TEXT NOT NULL REFERENCES sessions(id),
    test_name TEXT NOT NULL,
    test_path TEXT NOT NULL,
    outcome TEXT NOT NULL,
    duration_ms INTEGER NOT NULL,
    timestamp TEXT NOT NULL,
    commit_hash TEXT NOT NULL,
    branch TEXT NOT NULL,
    environment TEXT NOT NULL,
    retry_count INTEGER NOT NULL DEFAULT 0,
    error_message TEXT,
    stack_trace TEXT
);

-- Computed flakiness scores (upserted)
CREATE TABLE flakiness_scores (
    test_name TEXT PRIMARY KEY,
    probability_flaky REAL NOT NULL,
    confidence REAL NOT NULL,
    pass_rate REAL NOT NULL,
    fail_rate REAL NOT NULL,
    total_runs INTEGER NOT NULL,
    consecutive_failures INTEGER NOT NULL DEFAULT 0,
    last_updated TEXT NOT NULL,
    bayesian_params TEXT NOT NULL
);

-- Quarantine entries
CREATE TABLE quarantine (
    test_name TEXT PRIMARY KEY,
    quarantined_at TEXT NOT NULL,
    reason TEXT NOT NULL,
    flakiness_score REAL NOT NULL DEFAULT 0.0,
    auto_quarantined INTEGER NOT NULL DEFAULT 0
);

Utility Functions

FunctionSignatureDescription
parse_timestampfn(s: &str) -> DateTime<Utc>Parses RFC 3339 timestamps, falls back to Utc::now()
duration_to_msfn(d: Duration) -> i64Converts Duration to milliseconds for storage
ms_to_durationfn(ms: i64) -> DurationConverts milliseconds back to Duration

See the Storage Reference for configuration and migration details.

Runner API

The runner module handles test discovery, execution, and result collection.

Test Discovery

discover_test_binaries

#![allow(unused)]
fn main() {
pub fn discover_test_binaries(
    project_root: &Path,
) -> Result<Vec<TestBinary>, NinetyNineError>
}

Discovers all test binaries in a Cargo project by running cargo test --no-run --message-format json-render-diagnostics and parsing the output.

Returns: A list of TestBinary structs, each containing the binary path, package name, and kind.

Errors: Returns BinaryDiscovery if cargo fails or output cannot be parsed.

list_tests_parallel

#![allow(unused)]
fn main() {
pub async fn list_tests_parallel(
    binaries: &[TestBinary],
    concurrency: usize,
) -> Result<Vec<TestCase>, NinetyNineError>
}

Lists all tests across multiple binaries concurrently. Each binary is invoked with --list --format terse and output lines ending with : test or : benchmark are parsed.

Uses tokio::sync::Semaphore for concurrency control.

Errors: Returns TestListing if a binary fails to produce test listings.

cargo_available

#![allow(unused)]
fn main() {
pub fn cargo_available() -> bool
}

Returns whether cargo is on PATH. The native runner builds test binaries via cargo test --no-run and executes them directly, so cargo is the only external tool required.

Test Execution

Executor

#![allow(unused)]
fn main() {
pub struct Executor<'a> {
    config: &'a ExecutionConfig,
}
}

Runs individual test cases with retry support.

Constructor:

#![allow(unused)]
fn main() {
pub fn new(config: &'a ExecutionConfig) -> Self
}

run_single

#![allow(unused)]
fn main() {
pub fn run_single(
    &self,
    test_case: &TestCase,
) -> Result<TestResult, NinetyNineError>
}

Executes a single test case with retries. Spawns the test binary with the --exact flag targeting the specific test.

Retry behavior:

  • Retries up to config.retries times on failure
  • Stops immediately on first pass
  • Applies config.retry_delay between attempts

Timeout: Uses polling-based detection (50ms intervals). Kills the process when the deadline is exceeded, returning TestOutcome::Timeout.

Outcome classification:

ConditionOutcome
Exit code 0Passed
panicked at in stderr/stdoutPanic
Deadline exceededTimeout
Other non-zero exitFailed

ExecutionConfig

#![allow(unused)]
fn main() {
pub struct ExecutionConfig {
    pub concurrency: usize,
    pub timeout: Duration,
    pub retries: u32,
    pub retry_delay: Duration,
}
}
FieldDefaultDescription
concurrency—Maximum parallel test binary invocations
timeout300sPer-test execution timeout
retries0Number of retry attempts on failure
retry_delay100msDelay between retry attempts

TestResult

#![allow(unused)]
fn main() {
pub struct TestResult {
    pub test_case: TestCase,
    pub outcome: TestOutcome,
    pub duration: Duration,
    pub stdout: String,
    pub stderr: String,
    pub attempt: u32,
}
}

Test Case Types

TestCase

#![allow(unused)]
fn main() {
pub struct TestCase {
    pub name: TestName,
    pub binary_path: PathBuf,
    pub binary_name: String,
    pub package_name: String,
    pub binary_kind: BinaryKind,
    pub kind: TestKind,
}
}

TestKind

#![allow(unused)]
fn main() {
pub enum TestKind {
    Test,
    Benchmark,
}
}

BinaryKind

#![allow(unused)]
fn main() {
pub enum BinaryKind {
    Lib,
    Bin,
    Test,
    Example,
}
}

Derived from Cargo metadata target kinds.

TestBinary

#![allow(unused)]
fn main() {
pub struct TestBinary {
    pub path: PathBuf,
    pub package_name: String,
    pub binary_name: String,
    pub kind: BinaryKind,
}
}

High-Level Runner

NativeRunner

#![allow(unused)]
fn main() {
pub struct NativeRunner { /* private fields */ }
}

Constructor:

#![allow(unused)]
fn main() {
pub fn new(project_root: &Path, config: ExecutionConfig) -> Self
}

Methods:

MethodDescription
discover_tests(&self, filter: &str)Discovers test cases, optionally filtered by name substring

RunnerBackend

#![allow(unused)]
fn main() {
pub enum RunnerBackend {
    Native(NativeRunner),
}
}

Extensible enum wrapping runner implementations. Currently supports native Cargo test execution.

Methods: native(), execution_config(), discover_tests() — all delegate to the inner NativeRunner.

Standalone Function

execute_iterations

#![allow(unused)]
fn main() {
pub fn execute_iterations(
    test_case: &TestCase,
    iterations: u32,
    config: &ExecutionConfig,
    environment: &TestEnvironment,
) -> Result<Vec<TestRun>, NinetyNineError>
}

Convenience function that runs a test for N iterations and converts results to TestRun records. Used by the main command handler.

Configuration API

The configuration module handles loading, parsing, and serializing the .ninety-nine.toml configuration file.

Loading

load_config

#![allow(unused)]
fn main() {
pub fn load_config(project_root: &Path) -> Result<Config, NinetyNineError>
}

Searches for .ninety-nine.toml in the given project root directory. Returns Config::default() if the file does not exist.

Errors: Returns ConfigParse if the file exists but contains invalid TOML, or ConfigIo if the file cannot be read.

default_config_toml

#![allow(unused)]
fn main() {
pub fn default_config_toml() -> Result<String, NinetyNineError>
}

Serializes Config::default() to pretty-printed TOML. Used by the init command.

backoff_base_delay

#![allow(unused)]
fn main() {
pub fn backoff_base_delay(strategy: &BackoffStrategy) -> Duration
}

Extracts the initial delay from a backoff strategy.

StrategyBase Delay
None0ms
Linear { delay_ms }delay_ms
Exponential { base_ms, .. }base_ms
Fibonacci { start_ms, .. }start_ms

Config Model

Config

Top-level configuration struct. All fields have sensible defaults.

#![allow(unused)]
fn main() {
pub struct Config {
    pub detection: DetectionConfig,
    pub retry: RetryConfig,
    pub quarantine: QuarantineConfig,
    pub storage: StorageConfig,
    pub reporting: ReportingConfig,
}
}

DetectionConfig

Controls flakiness detection behavior.

#![allow(unused)]
fn main() {
pub struct DetectionConfig {
    pub min_runs: u32,
    pub confidence_threshold: f64,
    pub window_size: u32,
    pub parallel_runs: u32,
    pub duration_regression: Option<DurationRegressionConfig>,
}
}
FieldDefaultDescription
min_runs10Minimum iterations per test
confidence_threshold0.95Statistical confidence required to classify as flaky
window_size100Maximum historical runs to consider
parallel_runs3Number of concurrent test executions
duration_regressionNoneDuration regression tuning; None applies the defaults (enabled, 10-run history, 2 standard deviations)

RetryConfig

Controls test retry behavior on failure.

#![allow(unused)]
fn main() {
pub struct RetryConfig {
    pub unit_test_retries: u32,
    pub backoff_strategy: BackoffStrategy,
    pub max_retry_time_secs: u64,
}
}
FieldDefaultDescription
unit_test_retries2Maximum retries per failing test; every attempt is recorded as its own run
backoff_strategyExponential(100ms, 2.0x, 5000ms)Delay strategy between retries
max_retry_time_secs300Time limit for a single test execution (seconds)

BackoffStrategy

#![allow(unused)]
fn main() {
pub enum BackoffStrategy {
    None,
    Linear { delay_ms: u64 },
    Exponential { base_ms: u64, factor: f64, max_ms: u64 },
    Fibonacci { start_ms: u64, max_ms: u64 },
}
}

QuarantineConfig

Controls automatic and manual test quarantine.

#![allow(unused)]
fn main() {
pub struct QuarantineConfig {
    pub enabled: bool,
    pub auto_quarantine: bool,
    pub threshold: QuarantineThreshold,
}
}
FieldDefaultDescription
enabledtrueEnable quarantine system
auto_quarantinefalseAutomatically quarantine tests exceeding thresholds

QuarantineThreshold

#![allow(unused)]
fn main() {
pub struct QuarantineThreshold {
    pub consecutive_failures: u32,
    pub failure_rate: f64,
    pub flakiness_score: f64,
}
}
FieldDefaultDescription
consecutive_failures3Consecutive failures before quarantine
failure_rate0.20Failure rate threshold (20%)
flakiness_score0.15Bayesian score threshold

StorageConfig

#![allow(unused)]
fn main() {
pub struct StorageConfig {
    pub backend: StorageBackendType,
    pub retention_days: u32,
    pub sqlite: Option<SqliteConfig>,
    pub postgres: Option<PostgresConfig>,
}
}
FieldDefaultDescription
backendSqliteStorage backend to use
retention_days90Days to retain test run data

StorageBackendType

#![allow(unused)]
fn main() {
pub enum StorageBackendType {
    Sqlite,
    Postgres,
}
}

SqliteConfig

#![allow(unused)]
fn main() {
pub struct SqliteConfig {
    pub database_path: PathBuf,
}
}

Default path: $XDG_DATA_HOME/ninety-nine/ninety-nine.db

PostgresConfig

#![allow(unused)]
fn main() {
pub struct PostgresConfig {
    pub connection_string: String,
    pub pool_size: u32,
}
}

ReportingConfig

#![allow(unused)]
fn main() {
pub struct ReportingConfig {
    pub console: ConsoleOutputConfig,
}

pub struct ConsoleOutputConfig {
    pub summary_only: bool,  // default: false
}
}

DurationRegressionConfig

#![allow(unused)]
fn main() {
pub struct DurationRegressionConfig {
    pub enabled: bool,
    pub min_history_runs: u32,
    pub threshold: DurationThreshold,
}

pub enum DurationThreshold {
    Multiplier(f64),
    StdDev(f64),
}
}

See the Configuration Reference for TOML examples.

Error Types

All fallible operations in cargo-ninety-nine return Result<T, NinetyNineError>.

NinetyNineError

A comprehensive error enum covering all failure modes, derived with thiserror.

#![allow(unused)]
fn main() {
pub enum NinetyNineError {
    ConfigParse { source: toml::de::Error },
    ConfigIo { path: PathBuf, source: std::io::Error },
    NoRunnerAvailable,
    RunnerExecution { message: String },
    InvalidConfig { message: String },
    BinaryDiscovery { message: String },
    TestListing { binary: PathBuf, message: String },
    TestNotFound { name: String },
    FilterParse { message: String },
    Io { source: std::io::Error },
    Json { source: serde_json::Error },
    Storage { source: rusqlite::Error },
    PostgresStorage { source: tokio_postgres::Error },
    PostgresPool { message: String },
}
}

Error Categories

Configuration Errors

VariantDisplay MessageCause
ConfigParsefailed to parse config: {source}TOML syntax error or invalid field value
ConfigIofailed to read config file {path}: {source}File exists but cannot be read (permissions, etc.)
InvalidConfiginvalid configuration: {message}Logical configuration error (e.g., Postgres backend without connection config)

Runner Errors

VariantDisplay MessageCause
NoRunnerAvailableno test runner available: install cargo-nextest or use cargo testNeither cargo-nextest nor cargo test found
RunnerExecutionrunner execution failed: {message}Test binary failed to spawn or produced unexpected output
BinaryDiscoverybinary discovery failed: {message}cargo test --no-run failed or produced unparseable output
TestListingtest listing failed for {binary}: {message}Test binary --list invocation failed
TestNotFoundtest not found: {name}Requested test does not exist in the project

Filter Errors

VariantDisplay MessageCause
FilterParsefilter parse error: {message}Invalid filter DSL syntax

Storage Errors

VariantDisplay MessageCause
Storagestorage error: {source}SQLite operation failed (auto-converted from rusqlite::Error)
PostgresStoragepostgres storage error: {source}PostgreSQL operation failed (auto-converted from tokio_postgres::Error)
PostgresPoolpostgres pool error: {message}Connection pool exhausted or configuration error

I/O and Serialization Errors

VariantDisplay MessageCause
Ioio error: {source}Generic I/O failure (auto-converted from std::io::Error)
Jsonjson serialization error: {source}JSON serialization/deserialization failure (auto-converted from serde_json::Error)

Automatic Conversions

The following From implementations allow ? propagation:

Source TypeTarget Variant
std::io::ErrorIo
serde_json::ErrorJson
rusqlite::ErrorStorage
tokio_postgres::ErrorPostgresStorage

Error Handling in Practice

All errors are displayed to stderr in the main entry point and cause the process to exit with code 1:

#![allow(unused)]
fn main() {
// Simplified from main.rs
match run(args).await {
    Ok(()) => {}
    Err(e) => {
        eprintln!("error: {e}");
        std::process::exit(1);
    }
}
}

When --verbose is enabled, the tracing subscriber is set to debug level, providing additional context before errors surface.

Architecture

Module Overview

cargo-ninety-nine
src/
  main.rs              Entry point, command dispatch
  lib.rs               Module re-exports
  env.rs               Git info, environment, CI provider detection
  orchestrator.rs      Test execution pipeline, session lifecycle, auto-quarantine
  error.rs             NinetyNineError enum (thiserror)
  analysis/
    mod.rs             Shared failure_rate() helper
    duration.rs        Duration regression detection
    pattern.rs         Failure pattern detection (time-of-day, environmental)
    trend.rs           Trend calculation (improving/stable/degrading)
  ci/
    mod.rs             CI re-exports
    workflow.rs        GitHub Actions / GitLab CI YAML generation
  cli/
    mod.rs             CLI argument definitions (clap derive)
    output.rs          Console/JSON output formatting
    export.rs          JUnit XML, HTML, CSV, JSON export
  config/
    mod.rs             Config loading from TOML
    model.rs           Config struct definitions with defaults
  detector/
    mod.rs             Detector re-exports
    bayesian.rs        Bayesian flakiness scoring
  filter/
    mod.rs             compile_filter(), build_eval_context()
    ast.rs             FilterExpr and Predicate AST nodes
    lexer.rs           Tokenizer for filter DSL
    parser.rs          Recursive descent parser
    eval.rs            Evaluator matching FilterExpr against TestMetadata
  runner/
    mod.rs             NativeRunner, RunnerBackend, execute_iterations()
    binary.rs          Test binary discovery via cargo_metadata
    listing.rs         Test case enumeration from binaries
    executor.rs        Per-test execution with timeout/retry
    detection.rs       Runner availability check
  storage/
    mod.rs             Storage trait (async), open_storage() factory
    backend.rs         StorageBackend enum with dispatch! macro
    mapping.rs         Shared row-to-domain-type conversion (RawTestRunRow, RawScoreRow)
    sqlite.rs          SQLite implementation (rusqlite, WAL mode)
    postgres.rs        PostgreSQL implementation (deadpool-postgres)
    schema.rs          SQL migration definitions
  tui/
    mod.rs             TUI entry points, event loop, terminal guard, signal handlers
    app.rs             Application state (ScoresApp, HistoryApp, TableState, SortField)
    input.rs           Key event mapping to actions
    render.rs          Ratatui widget rendering (scores table, detail overlay, history)
  types/
    mod.rs             Type re-exports
    test_run.rs        TestRun, TestOutcome, TestEnvironment
    test_name.rs       TestName newtype
    flakiness.rs       FlakinessScore, BayesianParams, FlakinessCategory
    trend.rs           TrendDirection, TrendSummary
    session.rs         RunSession, ActiveSession, QuarantineEntry
    analysis.rs        FailurePattern, PatternType

Test Execution Flow

CLI (clap)
    |
    v
Config Loader ---------> .ninety-nine.toml
    |
    v
Filter DSL (optional)
  lexer --> parser --> FilterExpr AST
    |
    v
Runner Backend (NativeRunner)
  +-------------------+
  | Binary Discovery  |  cargo test --no-run --message-format json
  |   (cargo_metadata)|
  +---------+---------+
            |
            v
  +-------------------+
  | Test Listing      |  binary --list --format terse
  |   (parallel via   |  (semaphore-bounded concurrency)
  |    tokio spawn)   |
  +---------+---------+
            |
            v
  +-------------------+
  | Filter Evaluation |  FilterExpr evaluated against TestMetadata
  |   (if DSL given)  |  (loads flaky/quarantined sets from storage)
  +---------+---------+
            |
            v
  +-------------------+
  | Executor          |  binary --exact test_name --nocapture
  |   (N iterations,  |  (parallel via spawn_blocking + Semaphore)
  |    timeout, retry) |
  +-------------------+
            |
            | Vec<TestRun>
            v
  +-------------------+
  | Bayesian Detector |  Beta(alpha, beta) posterior
  |   + Analysis      |  --> FlakinessScore
  |   + Duration      |  --> DurationRegression (optional)
  +-------------------+
            |
      +-----+-----+-------+
      v           v       v
  Storage      Reporter  TUI
  (SQLite/     (console/  (ratatui,
   Postgres)    JSON/      interactive
                export)    scores/history)

Filter DSL Pipeline

The filter system has three stages:

  1. Lexer (filter/lexer.rs) – tokenizes the input string into Token values: LParen, RParen, And (&), Or (|), Not (!), and Ident(String).

  2. Parser (filter/parser.rs) – recursive descent parser that builds a FilterExpr AST. Supports binary operators (&, |), unary negation (!), parenthesized grouping, and function-call predicates like test(pattern), package(name), binary(name), kind(lib|bin|test|example). Bare identifiers are resolved as keywords (flaky, quarantined, all) or treated as test name regex patterns.

  3. Evaluator (filter/eval.rs) – evaluates a FilterExpr against a TestMetadata struct containing the test name, package name, binary name, and binary kind. An EvalContext holds pre-loaded sets of flaky and quarantined test names from storage.

Storage Abstraction

The Storage trait defines 13 async methods for persisting and querying test data. Two backends implement it:

  • SqliteStorage – uses rusqlite with WAL mode and Mutex<Connection> for thread safety. Synchronous operations run inside async method signatures. Migrations use PRAGMA user_version.

  • PostgresStorage – uses deadpool-postgres for connection pooling with configurable pool size and timeouts. Migrations use a schema_migrations table.

The StorageBackend enum wraps both backends and dispatches calls via a dispatch! macro, keeping the orchestration layer backend-agnostic.

#![allow(unused)]
fn main() {
pub enum StorageBackend {
    Sqlite(SqliteStorage),
    Postgres(PostgresStorage),
}
}

The open_storage() factory function reads the config to determine which backend to initialize.

Type System

TestName

A newtype wrapper around String that prevents confusion with other string fields (branch names, commit hashes, error messages). Implements Deref<Target=str>, AsRef<str>, Display, From<String>, From<&str>, and PartialEq<str>. Used in TestRun, FlakinessScore, QuarantineEntry, TrendSummary, DurationRegression, and TestCase.

ActiveSession

Session lifecycle type. ActiveSession::start() creates a running session, and to_run_session() produces the storable RunSession snapshot without consuming it; the orchestrator finishes the session by id once the suite completes.

FlakinessScore and FlakinessCategory

FlakinessScore holds the Bayesian posterior parameters (alpha, beta, posterior mean/variance, credible interval) alongside aggregate statistics (pass rate, fail rate, consecutive failures). FlakinessCategory classifies scores into five levels:

Score RangeCategory
< 0.01Stable
0.01 - 0.05Occasional
0.05 - 0.15Moderate
0.15 - 0.30Frequent
>= 0.30Critical

Key Design Decisions

Native Test Runner

Instead of wrapping cargo test as a subprocess for the full run, the tool uses a three-layer native pipeline:

  1. Binary discovery – uses cargo_metadata to parse cargo test --no-run JSON output, extracting test binary paths with package name and binary kind.
  2. Test listing – executes each binary with --list --format terse to enumerate individual test names. Runs in parallel across binaries using tokio::spawn_blocking with semaphore-bounded concurrency.
  3. Per-test execution – runs each test individually via duct, using --exact for isolation. Supports configurable timeout, retry count, and retry delay.

This gives precise per-test timing, retry control, and outcome classification without parsing human-readable test output.

Parallel Execution

Test execution uses tokio::spawn_blocking with a Semaphore to bound concurrency. The semaphore limit comes from detection.parallel_runs in the config (default: 3). Each test iteration acquires a permit before spawning a blocking task.

Subprocess Management

Test execution uses duct for subprocess control with:

  • Poll-based timeout – try_wait() in a loop with a 50ms polling interval, kill() on deadline
  • Output capture – stdout and stderr captured for failure analysis
  • Unchecked mode – non-zero exit codes are handled as test outcomes, not errors

Storage Backend Dispatch

The StorageBackend enum uses a dispatch! macro to forward all 13 trait methods to the underlying backend. This avoids 130+ lines of boilerplate match arms while keeping the dispatch zero-cost.

Error Handling

All errors flow through NinetyNineError (thiserror), with variants for each subsystem: config, storage, binary discovery, test listing, runner execution, filter parsing, Postgres pool, and I/O. The ? operator propagates errors to the top-level handler.

Bayesian Detection

Overview

cargo ninety-nine uses Bayesian inference to estimate the probability that a test is flaky. This approach provides calibrated uncertainty estimates rather than simple pass/fail ratios.

The Model

Prior

The model starts with a uniform (uninformative) prior: Beta(1, 1). This represents no prior knowledge — any flakiness probability from 0 to 1 is equally likely before observing data.

Posterior Update

After observing test runs, the posterior distribution is:

Beta(alpha, beta)

where:
  alpha = prior_alpha + failures
  beta  = prior_beta  + passes

The posterior mean is used as the flakiness probability:

P(flaky) = alpha / (alpha + beta)

Credible Interval

A 95% credible interval is computed from the Beta distribution’s inverse CDF:

CI = [Beta.inverse_cdf(0.025), Beta.inverse_cdf(0.975)]

This interval narrows as more data is collected.

Confidence

Confidence is derived from the width of the credible interval:

confidence = 1.0 - (CI_upper - CI_lower)

Narrow intervals yield high confidence; wide intervals yield low confidence.

Classification

A test is classified as flaky when:

P(flaky) > 0.01 AND confidence >= confidence_threshold

The confidence_threshold is configurable (default: 0.95).

Flakiness Categories

CategoryP(flaky) Range
Stable< 1%
Occasional1% — 5%
Moderate5% — 15%
Frequent15% — 30%
Critical> 30%

Practical Implications

Number of Runs

With the Beta(1,1) prior:

  • 10 runs, 1 failure: P(flaky) = 2/12 = 16.7%, wide CI → low confidence
  • 100 runs, 1 failure: P(flaky) = 2/102 = 2.0%, narrow CI → high confidence
  • 100 runs, 0 failures: P(flaky) = 1/102 = 1.0%, classified as Stable

More runs yield narrower credible intervals and more reliable classifications.

Prior Effect

The uniform prior has minimal effect when there are many observations. With 100+ runs, the prior contributes < 2% to the posterior. For small sample sizes (< 10 runs), the prior pulls estimates toward 50%, which is conservative.

Stored Parameters

Each FlakinessScore record includes the full Bayesian parameters:

FieldDescription
alphaPosterior alpha (prior + failures)
betaPosterior beta (prior + passes)
posterior_meanalpha / (alpha + beta)
posterior_variance(alpha * beta) / (total^2 * (total + 1))
credible_interval_lower2.5th percentile of posterior
credible_interval_upper97.5th percentile of posterior

Analysis and Patterns

Pattern Detection

After running tests, cargo ninety-nine analyzes failure data to identify patterns that might explain why tests are flaky.

Pattern Types

PatternDetection Method
TimeOfDayFailures concentrated at a specific hour (3x expected rate)
EnvironmentalFailure rate differs > 15% between CI and local environments
RandomFailures present but no discernible pattern (fallback)

Time-of-Day Detection

Requires at least 5 failures. Counts failures by hour (0-23 UTC) and computes:

concentration = max_hour_count / expected_per_hour

If concentration >= 3.0, a TimeOfDay pattern is reported with:

  • correlation: min(concentration - 1.0, 1.0)
  • examples: the peak hour identified

This pattern suggests timing-dependent failures (e.g., midnight log rotation, scheduled background jobs).

Environmental Detection

Requires at least 3 CI runs and 3 local runs. Compares failure rates between environments:

diff = |ci_fail_rate - local_fail_rate|

If diff >= 0.15, an Environmental pattern is reported, identifying which environment has the higher failure rate. This pattern suggests resource-dependent failures (e.g., CPU count, memory limits, file system speed).

Random Pattern

If failures exist but no specific pattern is detected, a Random pattern is reported. This is the fallback – it means failures occur but do not correlate with time or environment.

Trend Analysis

Trend analysis compares recent failure rates to previous failure rates within a sliding window.

How It Works

  1. Filter runs for a specific test name
  2. Take up to window_size most recent runs (configurable, default: 100)
  3. Split into two halves: recent (first half) and previous (second half)
  4. Compute failure rate for each half
  5. Classify the delta:
DeltaDirection
> 5%Degrading
< -5%Improving
within +/- 5%Stable

Requirements

At least 4 runs are required for trend calculation. Fewer runs return None.

  • cargo ninety-nine status <test_name> – shows trend direction and delta
  • After test – degrading trends are highlighted in the post-detection analysis

Duration Regression Analysis

Duration regression detection identifies tests whose execution time has significantly increased compared to their historical average. This helps catch performance regressions early, even when tests still pass.

How It Works

  1. Collect all run durations for a test (most recent first)
  2. Require at least min_history_runs total data points
  3. Separate the latest run from historical runs
  4. Compute the mean and sample standard deviation of historical durations only (the latest run is excluded to avoid polluting the baseline with the spike being measured)
  5. Apply a standard deviation floor of 1% of the mean (handles identical-duration histories where raw std_dev is zero)
  6. Compute the z-score:
deviation = (latest_duration - historical_mean) / effective_std_dev
  1. If deviation > threshold, a regression is reported with the actual slowdown ratio (latest / mean)

Configuration

Duration regression detection runs by default with a 10-run history requirement and a 2 standard-deviation threshold. Provide the section in .ninety-nine.toml to tune those values, or set enabled = false to switch the check off:

[detection.duration_regression]
enabled = true
min_history_runs = 10

# Flag when latest run exceeds 2.5 standard deviations above the mean
[detection.duration_regression.threshold]
StdDev = 2.5

Two threshold variants are available:

VariantDescriptionExample
StdDev(f64)Number of standard deviations above the meanStdDev = 2.0 flags tests running >2 sigma slower
Multiplier(f64)Multiple of the historical meanMultiplier = 3.0 flags tests taking >3x their average

Output

When a regression is detected, the CLI displays:

  SLOW tests::my_test — 500ms (mean: 100ms, 5.0x)

The DurationRegression struct contains:

FieldDescription
test_nameFully qualified name of the affected test
current_msDuration of the latest run in milliseconds
mean_msHistorical mean duration in milliseconds (excludes latest)
std_dev_msStandard deviation of historical durations
deviation_factorHow many standard deviations above the historical mean

The display shows the actual slowdown ratio (current_ms / mean_ms) rather than the z-score.

Edge Cases

  • If the effective standard deviation is near zero after flooring, no regression is reported
  • If fewer than min_history_runs total runs exist, no regression check is performed
  • If only one historical run exists (after separating the latest), the mean is that single value and the floor applies

Storage

Storage Trait

The Storage trait defines the async interface for all data persistence. Both backends implement all 13 methods:

#![allow(unused)]
fn main() {
pub trait Storage: Send + Sync {
    async fn store_session(&self, session: &RunSession) -> Result<(), NinetyNineError>;
    async fn finish_session(&self, session_id: &Uuid, test_count: u32, flaky_count: u32) -> Result<(), NinetyNineError>;
    async fn store_test_run(&self, run: &TestRun, session_id: &Uuid) -> Result<(), NinetyNineError>;
    async fn store_flakiness_score(&self, score: &FlakinessScore) -> Result<(), NinetyNineError>;
    async fn get_test_runs(&self, test_name: &str, limit: u32) -> Result<Vec<TestRun>, NinetyNineError>;
    async fn get_recent_sessions(&self, limit: u32) -> Result<Vec<RunSession>, NinetyNineError>;
    async fn get_all_scores(&self) -> Result<Vec<FlakinessScore>, NinetyNineError>;
    async fn get_score(&self, test_name: &str) -> Result<Option<FlakinessScore>, NinetyNineError>;
    async fn quarantine_test(&self, test_name: &str, reason: &str, score: f64, auto: bool) -> Result<(), NinetyNineError>;
    async fn unquarantine_test(&self, test_name: &str) -> Result<(), NinetyNineError>;
    async fn get_quarantined_tests(&self) -> Result<Vec<QuarantineEntry>, NinetyNineError>;
    async fn is_quarantined(&self, test_name: &str) -> Result<bool, NinetyNineError>;
    async fn purge_older_than(&self, days: u32) -> Result<u64, NinetyNineError>;
}
}

The StorageBackend enum wraps both backends and dispatches calls via a dispatch! macro:

#![allow(unused)]
fn main() {
pub enum StorageBackend {
    Sqlite(SqliteStorage),
    Postgres(PostgresStorage),
}
}

The open_storage() factory function reads the config to initialize the correct backend.

SQLite Backend (Default)

SQLite is the default backend, using rusqlite with the bundled SQLite library. No external database is required.

Default Location

$XDG_DATA_HOME/ninety-nine/ninety-nine.db

On Linux this is typically ~/.local/share/ninety-nine/ninety-nine.db. The parent directory is created automatically if it does not exist.

Features

  • WAL mode – enabled on open for concurrent read access during detection
  • Foreign keys – enforced via PRAGMA foreign_keys=ON
  • Thread safety – Mutex<Connection> guards all access; async trait methods run synchronously within async signatures
  • Bundled SQLite – no system SQLite dependency required

Configuration

[storage]
backend = "Sqlite"
retention_days = 90

[storage.sqlite]
database_path = "/custom/path/ninety-nine.db"

If storage.sqlite is omitted, the default path is used.

Migrations

Schema migrations use SQLite’s PRAGMA user_version. Each migration increments the version. Migrations are idempotent – running the tool against an already-migrated database is safe.

PostgreSQL Backend

PostgreSQL support uses deadpool-postgres for connection pooling with tokio-postgres for async queries.

Features

  • Connection pooling – configurable pool size via deadpool-postgres
  • Pool timeouts – 30s wait, 10s create, 5s recycle
  • Native async – all queries use async tokio-postgres directly
  • Schema migrations – tracked via a schema_migrations table with version and timestamp

Configuration

[storage]
backend = "Postgres"
retention_days = 90

[storage.postgres]
connection_string = "postgresql://user:password@localhost:5432/ninety_nine"
pool_size = 8

Note: Selecting backend = "Postgres" without providing [storage.postgres] will result in a configuration error at startup.

Migrations

PostgreSQL migrations use a schema_migrations table instead of pragmas. The migration system tracks applied versions and only runs new migrations. The schema is identical to SQLite in structure.

Schema

The database has four tables, identical in structure across both backends.

run_sessions

Tracks each detection run session.

ColumnTypeDescription
idTEXT (UUID)Session identifier
started_atTEXT/TIMESTAMPTZWhen the session started
finished_atTEXT/TIMESTAMPTZWhen the session finished (nullable)
test_countINTEGERTotal tests analyzed
flaky_countINTEGERTests classified as flaky
commit_hashTEXTGit commit at time of run
branchTEXTGit branch at time of run

test_runs

Individual test execution results.

ColumnTypeDescription
idTEXT (UUID)Run identifier
session_idTEXT (FK)Parent session
test_nameTEXTFully qualified test name
test_pathTEXTBinary path
outcomeTEXTpassed, failed, timeout, panic, ignored
duration_msINTEGER/BIGINTExecution time in milliseconds
timestampTEXT/TIMESTAMPTZWhen this run occurred
commit_hashTEXTGit commit
branchTEXTGit branch
retry_countINTEGERNumber of retries used
error_messageTEXTStderr/stdout on failure (nullable)
stack_traceTEXTStack trace if available (nullable)
env_osTEXTOperating system
env_rust_versionTEXTRust toolchain version
env_cpu_countINTEGERCPU core count
env_memory_gbREAL/DOUBLE PRECISIONSystem memory in GB
env_is_ciINTEGER/BOOLEANWhether running in CI
env_ci_providerTEXTCI provider name (nullable)

Indexed on test_name, timestamp, and session_id.

flakiness_scores

Latest computed flakiness scores, upserted on test_name.

ColumnTypeDescription
test_nameTEXT (PK)Fully qualified test name
probability_flakyREALBayesian P(flaky)
confidenceREAL1 - credible interval width
pass_rateREALPasses / total
fail_rateREALFailures / total
total_runsINTEGERNumber of runs
consecutive_failuresINTEGERTrailing failures
last_updatedTEXT/TIMESTAMPTZLast computation time
alphaREALBeta distribution alpha parameter
betaREALBeta distribution beta parameter
posterior_meanREALPosterior mean
posterior_varianceREALPosterior variance
ci_lowerREAL95% credible interval lower bound
ci_upperREAL95% credible interval upper bound

quarantine

Quarantined test records.

ColumnTypeDescription
test_nameTEXT (PK)Fully qualified test name
quarantined_atTEXT/TIMESTAMPTZWhen quarantined
reasonTEXTReason for quarantine
flakiness_scoreREALP(flaky) at time of quarantine
auto_quarantinedINTEGER/BOOLEANWhether auto-quarantined

Indexed on quarantined_at.

Data Retention

Old test runs are automatically purged based on the retention_days config (default: 90 days). Purging happens after each test session. Only the test_runs table is purged – flakiness_scores and quarantine entries persist.

[storage]
retention_days = 30

CLI Reference

Synopsis

cargo ninety-nine [OPTIONS] <COMMAND>

Global Options

OptionDefaultDescription
--project-dir <PATH>.Project root directory
--output <FORMAT>consoleOutput format: console or json
-N, --non-interactivefalseDisable TUI, use plain text output
-v, --verbosefalseEnable verbose output

By default, test, status, and history launch an interactive TUI. Use -N for CI pipelines, scripts, or piped output.


Commands

diagnose

Multi-phase stress → isolation → classify pipeline. See Multi-phase Diagnose.

cargo ninety-nine diagnose [OPTIONS] [FILTER_EXPR]
ArgumentDefaultDescription
[FILTER_EXPR]noneFilter expression (same DSL as test)
--stress-runs <N>config (3)Full-binary multi-threaded suite runs
--isolation-runs <N>config (10)Serial --exact runs per candidate
--recordfalseAttempt rr recording for Intrinsic failures
--no-recordfalseDisable recording even if config enables it
--record-dir <PATH>configDirectory for rr traces
--confidence <F>configBayesian confidence threshold

test

Run tests and detect flakiness. Each discovered test is executed multiple times, scored with Bayesian inference, and results are stored for trend analysis.

cargo ninety-nine test [OPTIONS] [FILTER_EXPR]
Argument/OptionDefaultDescription
[FILTER_EXPR]noneFilter expression (DSL or test name pattern)
-n, --iterations <N>from config (min_runs, default 10)Number of times to run each test
--confidence <FLOAT>from config (confidence_threshold, default 0.95)Confidence threshold for flaky classification

Examples:

# Run all tests 10 times each (defaults)
cargo ninety-nine test

# Run tests matching a substring
cargo ninety-nine test my_module

# Run with more iterations and stricter confidence
cargo ninety-nine test -n 25 --confidence 0.99

# Use filter DSL to run only flaky tests
cargo ninety-nine test "flaky"

# Combine filter predicates
cargo ninety-nine test "test(my_module) & !quarantined"

Filter DSL

The optional FILTER_EXPR argument accepts a domain-specific language for filtering tests. If the expression contains no DSL operators, it is treated as a plain test name substring filter.

Predicates:

PredicateDescription
test(pattern)Match test names by regex pattern
package(name)Match tests in a package (substring match)
binary(name)Match tests from a specific binary (substring match)
kind(type)Match by binary kind: lib, bin, test, example
flakyMatch tests previously detected as flaky
quarantinedMatch quarantined tests
allMatch all tests
bare_wordTreated as test(bare_word) regex pattern

Operators:

OperatorMeaning
&AND – both sides must match
|OR – either side must match
!NOT – negate the following expression
( )Grouping

Examples:

# Tests matching a regex
cargo ninety-nine test "test(my_module::.*)"

# Flaky tests that are not quarantined
cargo ninety-nine test "flaky & !quarantined"

# Tests in a specific package or binary kind
cargo ninety-nine test "package(my_crate) | kind(test)"

# Complex expression with grouping
cargo ninety-nine test "(flaky | test(slow)) & !quarantined"

init

Initialize a .ninety-nine.toml configuration file in the project root.

cargo ninety-nine init [OPTIONS]
OptionDescription
--forceOverwrite existing config file

status

Show flakiness status for tests. Without a test name, launches the interactive TUI (or prints a table with -N).

cargo ninety-nine status [TEST_NAME]
ArgumentDescription
[TEST_NAME]Show detailed status for a specific test. Omit to browse all scores interactively.

When a test name is provided, shows: category, P(flaky), pass rate, total runs, consecutive failures, credible interval, trend direction, failure patterns, and recent run history.


history

Show detection session history. Launches the interactive TUI by default (or prints a table with -N).

cargo ninety-nine history [OPTIONS] [FILTER]
Argument/OptionDefaultDescription
[FILTER]noneFilter by test name
-n, --limit <N>20Maximum sessions to show

export

Export flakiness data to a file.

cargo ninety-nine export <FORMAT> <PATH>
ArgumentValuesDescription
<FORMAT>junit, html, csv, jsonExport format
<PATH>file pathOutput file path

quarantine

Manage test quarantine.

quarantine list

cargo ninety-nine quarantine list

Lists all quarantined tests with their quarantine date, reason, flakiness score, and whether they were auto-quarantined.

quarantine add

cargo ninety-nine quarantine add <TEST_NAME> [OPTIONS]
OptionDefaultDescription
--reason <TEXT>"manually quarantined"Reason for quarantine

quarantine remove

cargo ninety-nine quarantine remove <TEST_NAME>

ci

CI integration helpers.

ci generate

Generate a CI workflow file for flaky test detection.

cargo ninety-nine ci generate <PROVIDER> [PATH]
ArgumentValuesDescription
<PROVIDER>github, gitlabCI provider
[PATH]file pathOutput file path (default: stdout)

Configuration Reference

All configuration lives in .ninety-nine.toml at the project root. Every field has a default value – you only need to specify overrides.

[detection]

Controls how tests are discovered and analyzed.

FieldTypeDefaultDescription
min_runsu3210Minimum number of runs per test
confidence_thresholdf640.95Confidence level required to classify as flaky
window_sizeu32100Number of recent runs used for trend analysis
parallel_runsu323Number of concurrent test executions

Detection uses Beta distribution posterior inference with conjugate priors. A test is only ever classified as flaky once at least one failure has actually been observed; long histories of clean passes always classify as stable.

[detection.duration_regression]

Optional configuration for detecting tests whose execution time has significantly increased.

FieldTypeDefaultDescription
enabledbool–Enable duration regression detection
min_history_runsu32–Minimum historical runs required before checking
thresholdenum–How to determine a regression (see below)

When the section is omitted entirely, the check runs with its long-standing defaults: a 10-run history requirement and a threshold of 2 standard deviations. Provide the section to tune those values, or set enabled = false to switch the check off:

[detection.duration_regression]
enabled = true
min_history_runs = 10

# Option A: flag when latest duration exceeds N standard deviations above the mean
[detection.duration_regression.threshold]
StdDev = 2.0

# Option B: flag when latest duration exceeds mean * multiplier
# [detection.duration_regression.threshold]
# Multiplier = 3.0

Threshold variants:

VariantDescription
StdDev(f64)Flag when the latest duration is more than N standard deviations above the historical mean
Multiplier(f64)Flag when the latest duration exceeds the historical mean multiplied by this factor

[diagnose]

Controls the multi-phase diagnose command (stress / isolation / optional rr).

FieldTypeDefaultDescription
stress_runsu323Full-binary multi-threaded suite iterations
isolation_runsu3210Serial isolation iterations per candidate
stress_threadsu320libtest threads (0 = host parallelism)
stress_timeout_secsu64300Timeout per stress binary iteration
recordboolfalseAttempt rr recording for Intrinsic failures
record_dirpath.ninety-nine/recordingsrr output directory
record_attemptsu3210Max rr attempts per Intrinsic candidate
chaosboolfalsePass --chaos to rr when recording

[detection] multi-phase

FieldTypeDefaultDescription
multi_phaseboolfalseWhen true, test runs diagnose for candidates then Bayesian suite for the rest

[quarantine.by_class]

FieldTypeDefaultDescription
intrinsicbooltrueAuto-quarantine Intrinsic diagnose classes
contentionboolfalseAuto-quarantine Contention classes
brokenbooltrueAuto-quarantine Broken classes

[retry]

Controls retry behavior when tests fail.

FieldTypeDefaultDescription
unit_test_retriesu322Number of retries for failed tests
backoff_strategyenumExponentialBackoff strategy between retries
max_retry_time_secsu64300Time limit for a single test execution (seconds)

Every attempt is recorded as its own run, so a test that fails and then passes on retry contributes both outcomes to its flakiness score — recovery on retry is itself flaky evidence, and raising the retry count therefore gathers more evidence per iteration rather than hiding failures.

Backoff Strategies

# No delay between retries
[retry]
backoff_strategy = "None"

# Fixed delay
[retry.backoff_strategy]
Linear = { delay_ms = 200 }

# Exponential backoff (default: base_ms=100, factor=2.0, max_ms=5000)
[retry.backoff_strategy]
Exponential = { base_ms = 100, factor = 2.0, max_ms = 5000 }

# Fibonacci sequence delays
[retry.backoff_strategy]
Fibonacci = { start_ms = 100, max_ms = 5000 }

[quarantine]

Controls test quarantine behavior.

FieldTypeDefaultDescription
enabledbooltrueEnable quarantine functionality
auto_quarantineboolfalseAutomatically quarantine tests exceeding thresholds

[quarantine.threshold]

Thresholds for auto-quarantine. A test must be classified as flaky AND exceed at least one threshold.

FieldTypeDefaultDescription
consecutive_failuresu323Consecutive trailing failures
failure_ratef640.20Overall failure rate
flakiness_scoref640.15Bayesian P(flaky)

[storage]

Controls data persistence. Two backends are supported: SQLite (default) and PostgreSQL.

FieldTypeDefaultDescription
backendenum"Sqlite"Storage backend: Sqlite or Postgres
retention_daysu3290Days to keep test run data; sessions left without any runs by the purge are removed with them

[storage.sqlite]

SQLite-specific settings. Only used when backend = "Sqlite".

FieldTypeDefaultDescription
database_pathpath$XDG_DATA_HOME/ninety-nine/ninety-nine.dbPath to the SQLite database file
[storage]
backend = "Sqlite"
retention_days = 90

[storage.sqlite]
database_path = ".ninety-nine/data.db"

[storage.postgres]

PostgreSQL-specific settings. Required when backend = "Postgres".

FieldTypeDefaultDescription
connection_stringstring–PostgreSQL connection URL
pool_sizeu32–Connection pool size
[storage]
backend = "Postgres"
retention_days = 90

[storage.postgres]
connection_string = "postgresql://user:password@localhost:5432/ninety_nine"
pool_size = 8

Warning: Setting backend = "Postgres" without providing [storage.postgres] will cause a startup error.

[reporting]

Controls output behavior.

[reporting.console]

FieldTypeDefaultDescription
summary_onlyboolfalseShow only summary counts instead of full table

Full Example

[detection]
min_runs = 20
confidence_threshold = 0.99
window_size = 50
parallel_runs = 4

[detection.duration_regression]
enabled = true
min_history_runs = 10

[detection.duration_regression.threshold]
StdDev = 2.5

[retry]
unit_test_retries = 3
max_retry_time_secs = 120

[retry.backoff_strategy]
Exponential = { base_ms = 200, factor = 2.0, max_ms = 10000 }

[quarantine]
enabled = true
auto_quarantine = true

[quarantine.threshold]
consecutive_failures = 5
failure_rate = 0.15
flakiness_score = 0.10

[storage]
backend = "Sqlite"
retention_days = 60

[storage.sqlite]
database_path = ".ninety-nine/data.db"

[reporting.console]
summary_only = false