Skip to main content
Each RL task is one real change from a private production codebase, rebuilt as a Harbor task that grades offline with a continuous reward. This page describes what a delivered task contains and how its numbers are produced.

How a task is made

How an RL task is made, in seven steps: a licensed repository, one real commit becomes one task, the before-state is rebuilt, a graded verifier is written, QC checks, calibration, and export. A task that fails a QC check goes back for repair. How an RL task is made, in seven steps: a licensed repository, one real commit becomes one task, the before-state is rebuilt, a graded verifier is written, QC checks, calibration, and export. A task that fails a QC check goes back for repair.
  1. A licensed repository. Private production code with its git history, cleaned of secrets and personal data before the pipeline reads it. It has passed the exposure check.
  2. One real commit becomes one task. Mechanical changes are dropped. Small neighboring commits in the same module can be combined into one task.
  3. The before-state is rebuilt as a Harbor task container at the parent commit. The solver’s workspace carries no git history, grading runs offline, and the original commit message is withheld.
  4. A graded verifier is written. Checks run the code instead of scanning it. Fakes are used only at external service boundaries.
  5. QC checks prove the task can be solved and resists gaming. See QC checks.
  6. Calibration. Solver runs place the task in a difficulty band.
  7. Export. The task ships as one bundle. A task that fails a QC check goes back for repair, or does not ship.

Bundle layout

The container’s workspace is /app. The solver runs as the unprivileged node user. The verifier wipes and restores its own test directory before every grade, so nothing the solver leaves behind can stand in for the graded suite.

task.toml

Reward

The reward is weighted checks passed divided by weighted total, with behavioral, integration, end-to-end, and runtime checks weighted 4, core and other checks 1, and setup, install, typecheck, build, and lint checks 0. On a 0 to 1 scale, doing nothing must score 0.1 or less, the reference solution 0.95 or more, and solver scores from 0.1 to 0.8 are deliverable. The reward is weighted checks passed divided by weighted total, with behavioral, integration, end-to-end, and runtime checks weighted 4, core and other checks 1, and setup, install, typecheck, build, and lint checks 0. On a 0 to 1 scale, doing nothing must score 0.1 or less, the reference solution 0.95 or more, and solver scores from 0.1 to 0.8 are deliverable. The reward is continuous from 0 to 1:
  • Core checks gate partial credit. Core checks are the few that prove the task was genuinely attempted. If any fails, the whole score is multiplied by the fraction of core checks that passed.
  • Synthesized checks carry a fixed share. Some tasks carry an added synthesized check layer. When it is present it is 50% of the reward, and the other layers make up the rest.
  • The verifier writes the reward to /logs/verifier/reward.txt. It reads results only from a private results file the solver cannot reach. Printed output is never scored.
A run scores 0 when:
  • the results file is missing or empty;
  • the verifier’s test tooling does not match its SHA-256 manifest;
  • the verifier leaves tracked files modified;
  • one grading pass runs past PD_GRADE_WALL_CAP_SEC.

QC checks

Every shipped task clears these checks:
  • The reference solution scores at least 0.95, run in a cold container.
  • Doing nothing scores at most 0.1, and no scored check passes on the untouched code.
  • The checks are audited for fairness against the instruction.
  • Mutation checks: renaming internal variables keeps the score at 0.95 or above, and removing one specified behavior lowers it.
  • Cheat probes score 0.1 or less: the reference solution changed to exit as soon as it loads, with and without printing forged test results.
  • Solver scores vary across runs, and the checks order solver attempts consistently by difficulty (Loevinger H of at least 0.40). Flat or incoherent tasks go back for repair.
  • Solver runs place the task in the deliverable band.
qa-evidence/ holds the reference, do-nothing, mutation, and cheat-probe runs. qc-report.md summarizes the baselines, the mutation results, and the calibration table.

Difficulty bands

Solver runs set each task’s band from its solver score: Only deliverable-band tasks count as tasks. A task whose solver runs score no better than doing nothing is treated as ungradeable, not hard, and goes back for repair. Frontier-model results from several model families can be run for a delivered set on request.

Checking overlap with your corpus

manifest.json carries a file_tree for the repository at the before-state, the code the solver starts from:
  • files maps every regular file to the SHA-256 of its contents. Paths are relative to the repository, /-separated, and Unicode NFC. .git and symbolic links are skipped.
  • root is the SHA-256 of the lines <path>\0<sha256>\n, one per file, sorted by the UTF-8 bytes of the path. The same paths and contents always give the same root.
To check overlap, hash the files in your corpus with SHA-256 and intersect the digests with the values of files. To compare a whole tree, compute its root the same way:

Delivery

Each task ships as a .tar.gz archive with its SHA-256 checksum. Check the archive before you unpack it, then run the reference solution to confirm the verifier in your own environment:
reproduction-kit.md lists the commands behind every number in the bundle, including model runs and the mutation checks. To run the cheat script, put cheat/solve.sh in place of solution/solve.sh in a copy of the task and run the oracle agent.