Guide

How CI outcomes become RLVR training data

Reinforcement learning from verifiable rewards (RLVR) trains a model on outcomes a program can check, not judgments a person makes. In Sediment's coding-agent context, the checkable outcome is a recorded CI result: an agent session produced a commit, CI ran, and the run passed or failed.

Sediment exports that evidence as task rows and trajectory rows. Task rows record the repository state, reference patch, and recorded results. Verification commands appear when an operator configures them. Trajectory rows preserve captured turns and their evidence. Both carry provenance that identifies the derivation policy.

Turn a CI run into a reward

A reward needs more than a green check. Sediment separates what CI reported from what it means, so every reward can be traced back to the exact recorded attempts. For the pipeline that produces the underlying evidence, read the methodology guide.

  1. Capture every attempt

    Sediment records each CI run attempt as a fact: the provider, the run and attempt identity, the provider's original result, and a standardized result. Capture doesn't parse logs or decide what a failure means.

  2. Resolve one verdict

    Resolution happens later, under defined rules. Attempts group by provider run, ordered by attempt number. The last passed or failed attempt supplies the workflow's verdict. Capture time and ingest order never select a verdict.

    A timeout, cancellation, skip, or unknown result stays recorded but supplies no verdict. Workflows that disagree on one commit produce no reward at all.

  3. Score reliability separately

    A run that flips between attempts — failure then pass, or pass then failure — keeps its verdict but resolves with suspected_flake: true and a default reliability of 0.0. Workflows that agree use the minimum reliability.

    The reward stays categorical. Reliability rides along so your pipeline can weight or filter rows; Sediment never converts it into a sample weight.

Read the fields of a task row

A task row answers six questions: what the task was, what patch solved it, how to check a solution, what CI recorded, what reward that supports, and how the row was derived. Field names below match the canonical sediment-rlvr-task v3 schema.

  • The task

    instance_id names the row: organization, session, and the first attributed commit's short hash in the verified repository. repo and base_commit give the repository and the checkout point the patch applies to — the parent of that repository's first attributed commit. problem_statement is the session's first user message, flattened to text.

  • The reference patch

    reference_patch is the recorded diff from base_commit to the exact verified commit — history your team wrote and CI checked. A trainer compares candidate patches against it. Sediment reads it from the repository mirror and never synthesizes it. Later commits without verifier evidence never extend it.

  • The rerun command

    verification.verification_command is the command that can run the checks again, such as pytest -q. An operator supplies it per repository in a configuration file. Sediment stores it for your training pipeline; it never derives the command from workflow YAML and never runs it itself. Without configuration, the field is omitted.

  • The recorded results

    verifier_results keeps the recorded CI outcomes for the selected commit exactly as captured, including results that carry no verdict. ci_resolution is the recomputed interpretation for the verified commit: the verdict, its reliability, whether a flake is suspected, and the outcome ids it used.

  • The reward

    reward_source is resolved_ci_pass or resolved_ci_fail. A session whose only results carry no verdict produces no task row rather than a guessed reward. The NeMo Gym target maps a pass to a numeric reward of 1.0 and a failure to 0.0.

  • The provenance

    provenance records the derivation policy version, the quarantine revision, and the policy digest. attribution_source says whether the session-to-commit link came from a Git note or a similarity estimate. session_commit_observation_ids lists the supporting observations when available. repository_identity names the provider, host, and immutable repository ID; it is null for unambiguous legacy evidence. split assigns the whole session to train or eval.

Inspect a worked example

This synthetic task row shows one complete chain: a session fixed a rounding bug, the fix reached commit b3d9e2f, and the CI workflow passed on its first attempt. The row uses legacy evidence without an immutable repository ID. It validates against the sediment-rlvr-task v3 schema. The row is indented for readability; exports use one JSON object per line.

{
  "instance_id": "example-org-session-01-b3d9e2f",
  "recipe_id": "rlvr_ci",
  "recipe_version": 1,
  "reward_source": "resolved_ci_pass",
  "repo": "example-org/payments",
  "base_commit": "7e1c0d92a4b85f36c9d01e7aab24c3f8d5e6a901",
  "problem_statement": "Round cash totals to the nearest cent before saving.",
  "reference_patch": "diff --git a/payments/cash.py b/payments/cash.py\n--- a/payments/cash.py\n+++ b/payments/cash.py\n@@ -1,2 +1,2 @@\n def total(amounts):\n-    return sum(amounts)\n+    return round(sum(amounts), 2)\n",
  "verification": {
    "verification_command": "pytest -q"
  },
  "verifier_results": [
    {
      "schema_version": 1,
      "repository_provider": null,
      "repository_host": null,
      "repository_id": null,
      "outcome_id": "3f7f0d6e-2b41-4a56-9d2c-8e5a1c0b7d43",
      "org_id": "example-org",
      "provider": "github_actions",
      "run_id": "1234567890",
      "run_attempt": 1,
      "repo": "example-org/payments",
      "commit_sha": "b3d9e2f81c6a45d70e92b3f1a8c4d5e6f7091a2b",
      "branch": "main",
      "result": "passed",
      "workflow_name": "CI",
      "workflow_id": "987",
      "workflow_path": ".github/workflows/ci.yml",
      "run_url": "https://github.com/example-org/payments/actions/runs/1234567890",
      "provider_result": "success",
      "error_type": null,
      "reason": null,
      "source_event_type": "github.workflow_run.completed",
      "source_spec_version": null,
      "source_event_id": null,
      "pr_number": null,
      "captured_at": "2026-08-14T09:30:00Z",
      "raw": {}
    }
  ],
  "ci_resolution": {
    "org_id": "example-org",
    "repo": "example-org/payments",
    "commit_sha": "b3d9e2f81c6a45d70e92b3f1a8c4d5e6f7091a2b",
    "verdict": "passed",
    "reliability": 1,
    "suspected_flake": false,
    "workflow_resolutions": [
      {
        "provider": "github_actions",
        "run_id": "1234567890",
        "workflow_id": "987",
        "workflow_name": "CI",
        "workflow_path": ".github/workflows/ci.yml",
        "branch": "main",
        "verdict": "passed",
        "reliability": 1,
        "suspected_flake": false,
        "source_outcome_ids": [
          "3f7f0d6e-2b41-4a56-9d2c-8e5a1c0b7d43"
        ],
        "verdict_outcome_id": "3f7f0d6e-2b41-4a56-9d2c-8e5a1c0b7d43",
        "non_verdict_outcome_ids": []
      }
    ],
    "source_outcome_ids": [
      "3f7f0d6e-2b41-4a56-9d2c-8e5a1c0b7d43"
    ],
    "verdict_outcome_ids": [
      "3f7f0d6e-2b41-4a56-9d2c-8e5a1c0b7d43"
    ],
    "non_verdict_outcome_ids": [],
    "provenance": {
      "policy_version": "3",
      "quarantine_revision": 0,
      "policy_digest": null
    },
    "repository_identity": null
  },
  "attribution_source": "git_notes",
  "split": "train",
  "provenance": {
    "policy_version": "3",
    "quarantine_revision": 0,
    "policy_digest": null
  },
  "session_commit_observation_ids": [],
  "repository_identity": null,
  "schema_id": "https://sediment.so/schemas/training-rows/sediment-rlvr-task/v3.json",
  "schema_version": 3
}

Read it bottom-up to audit the reward. reward_source points at ci_resolution, whose verdict_outcome_ids name the exact recorded outcome in verifier_results that supplied the verdict. One passing first attempt means reliability: 1.0 and no suspected flake. attribution_source: "git_notes" says the session-to-commit link was recorded, not estimated.

A first attempt that failed and a retry that passed would keep verdict: "passed" but carry reliability: 0.0 and suspected_flake: true — same reward, different trust.

Choose an export target

Run sediment export rlvr --target <target> with one of three targets. Every target is a stateless projection over the same recorded rollouts: it reads evidence and mirrors, and it never runs a verifier or invents missing fields.

  • sediment

    Writes tasks.jsonl and rollouts.jsonl, plus an optional experimental environment.yaml. This is the lossless audit projection: task rows retain repository and CI evidence; Rollout rows retain captured turns, tool calls, and decisions.

  • swe-bench

    Writes tasks.jsonl in the SWE-bench task shape, with Sediment evidence under metadata. The patch comes only from a recorded terminal pass. Fields Sediment's facts don't contain — issue ids, test patches, FAIL_TO_PASS lists — are omitted, not invented.

  • nemo-gym

    Writes rollouts.jsonl mapped to NeMo Gym's rollout boundary. A resolved pass becomes reward 1.0, a failure 0.0, and anything else omits the reward. The mapping names the boundary; it doesn't claim direct compatibility with any trainer release.

Trajectory rows need the full interaction sequence, which only gateway-routed model calls provide. Copilot's calls can't route through the gateway, so Copilot sessions contribute decisions and CI evidence but not trajectories.

Limitations

Review these four limits before you train on an RLVR export. Each affects how much a reward should be trusted, not just how it was produced.

  • Leakage

    The built-in split keeps each session on one side, but it doesn't isolate repositories, similar prompts, or external task identifiers across sessions. The reference patch is your own history, so near-identical fixes can appear on both sides.

    Before you evaluate, compare every exported row with your benchmark's manifest through its source identifiers, and exclude pilot or development tasks. The row-level split value alone doesn't prove a holdout.

  • Flaky verifiers

    A retry that flips the result keeps its verdict with reliability 0.0 and suspected_flake: true. The numeric reward stays exactly 1.0 or 0.0, so a reward can sit beside zero reliability. That pairing is intentional: honor the reliability field when you weight or filter, or a flaky pass trains as a clean one.

  • Reward integrity

    The reward reflects a recorded result, not a fresh run. Sediment never executes the rerun command, and repository or dependency drift can make a recorded result unreproducible later. The rerun command itself is operator-supplied and unverified.

    If Facts, policy, snapshot and quarantine revision, and mirror refs remain unchanged, repeated exports within one mode preserve JSONL bytes and row order. Compare parsed values across direct and bundle exports because their nested key order can differ. Use a fresh output directory for each export so generations can't mix. Validate rows against the schema before an adapter strips metadata.

  • CI is not semantic proof

    A pass means the recorded checks found no failure. Weak or missing tests pass incorrect code, and a failure doesn't prove the patch caused it. An RLVR reward is only as meaningful as the checks behind it, so invest in the test suite before you invest in the reward.

Where evidence is missing, Sediment skips the row and counts the reason under a closed vocabulary. The RLVR export guide lists every skip reason.

Next steps

  • RLVR export

    Run the export and read each target's full contract and skip reasons.

  • Run derivations

    Set policy, derive a reviewed bundle, and export from it.

  • Export docs

    Compare RLVR with SFT and DPO and audit a split before training.

  • Schemas

    Validate rows against the versioned JSON schemas.