LIVEdataset aec-bench@releasetasks 552models 18last submission · built
civilno-tool

Drainage Model Run Provenance Issue Review Package

Review a generated drainage-model evidence packet across governing input revisions, model manifest, run register, report, and downstream design memo; determine which state governs and whether the package is ready to issue.

no-tool: The model must reason numerically unaided.

How this task is generated

One template produces many comparable benchmark tasks while keeping the scoring contract fixed.

  1. 01

    Template

    The reusable contract shown on this page.

  2. 02

    Scenario

    An archetype and site context are sampled.

  3. 03

    Difficulty tier

    Inputs may be hidden at harder tiers.

  4. 04

    Task prompt

    The model responds with the declared outputs.

Parameters

Inputs the model receives, and the outputs it is scored on.

Inputs

10 inputs

Always given

Included directly in every task prompt.

4
  • Current catchment area

    current_catchment_area_ha

    Governing catchment area

    16 – 48 ha
  • Report peak flow

    report_peak_flow_m3_s

    Peak flow reported by the reviewed model run

    1.2 – 4.8 m3/s
  • Report max hgl

    report_max_hgl_m_ahd

    Maximum HGL reported by the reviewed model run

    11.2 – 24.8 m AHD
  • Maximum continuity error

    maximum_continuity_error_percent

    Maximum continuity error owned by the criteria memo

    1.5 – 2.5 percent

Hidden at higher difficulty

Visible in easier tasks and withheld in one or more harder tiers.

6
  • Previous area delta

    previous_area_delta_ha

    Hidden area delta between current and superseded catchment revisions

    Hidden at easy and medium and hard difficulty.

    0.4 – 2.2 ha
  • Continuity margin target

    continuity_margin_target_percent

    Hidden clean-case continuity margin

    Hidden at easy and medium and hard difficulty.

    0.2 – 0.8 percent
  • Continuity deficit

    continuity_deficit_percent

    Hidden failed-case continuity deficit

    Hidden at easy and medium and hard difficulty.

    0.15 – 0.6 percent
  • Prior peak flow delta

    prior_peak_flow_delta_m3_s

    Hidden difference between current and superseded report peak flow

    Hidden at easy and medium and hard difficulty.

    0.12 – 0.55 m3/s
  • Prior hgl delta

    prior_hgl_delta_m

    Hidden difference between current and superseded report HGL

    Hidden at easy and medium and hard difficulty.

    0.08 – 0.35 m
  • Packet variant

    packet_variant

    Hidden source-packet condition

    Hidden at easy and medium and hard difficulty.

    cleanmissing_manifest_catchment_revisionstale_catchment_revisionreport_run_id_mismatchcontinuity_limit_exceededscenario_copy_forwarddownstream_memo_stale_reportopen_critical_comment+1

Scored outputs

27 outputs

Prv 01 status

prv_01_status

PRV-01 status code

Scores if within ±1% of the reference value.

Prv 02 status

prv_02_status

PRV-02 status code

Scores if within ±1% of the reference value.

Prv 03 status

prv_03_status

PRV-03 status code

Scores if within ±1% of the reference value.

Prv 04 status

prv_04_status

PRV-04 status code

Scores if within ±1% of the reference value.

Prv 05 status

prv_05_status

PRV-05 status code

Scores if within ±1% of the reference value.

Prv 06 status

prv_06_status

PRV-06 status code

Scores if within ±1% of the reference value.

Prv 07 status

prv_07_status

PRV-07 status code

Scores if within ±1% of the reference value.

Prv 08 status

prv_08_status

PRV-08 status code

Scores if within ±1% of the reference value.

Prv 09 status

prv_09_status

PRV-09 status code

Scores if within ±1% of the reference value.

All input revisions match score

all_input_revisions_match_score

Whether the run manifest cites every governing input revision

Scores if within ±1% of the reference value.

Scenario match score

scenario_match_score

Whether the run manifest uses the governing design storm

Scores if within ±1% of the reference value.

Report run match score

report_run_match_score

Whether the report run ID matches the run register

Scores if within ±1% of the reference value.

Continuity error

continuity_error_percent

Reported continuity error

Scores if within ±2% of the reference value.

Continuity margin

continuity_margin_percent

Continuity-error acceptance margin

Scores if within ±2% of the reference value.

Report peak flow

report_peak_flow_m3_s

Peak flow in the reviewed report

Scores if within ±2% of the reference value.

Memo peak flow

memo_peak_flow_m3_s

Peak flow adopted by the design memo

Scores if within ±2% of the reference value.

Peak flow propagation delta

peak_flow_propagation_delta_m3_s

Absolute report-to-memo peak-flow difference

Scores if within ±2% of the reference value.

Report max hgl m ahd

report_max_hgl_m_ahd

Maximum HGL in the reviewed report

Scores if within ±2% of the reference value.

Memo max hgl m ahd

memo_max_hgl_m_ahd

Maximum HGL adopted by the design memo

Scores if within ±2% of the reference value.

Hgl propagation delta

hgl_propagation_delta_m

Absolute report-to-memo HGL difference

Scores if within ±2% of the reference value.

Run applicability code

run_applicability_code

Model-run applicability code

Scores if within ±1% of the reference value.

Report applicability code

report_applicability_code

Model-report applicability code

Scores if within ±1% of the reference value.

Design claim support code

design_claim_support_code

Downstream design-claim support code

Scores if within ±1% of the reference value.

Readiness code

readiness_code

Issue-readiness code

Scores if within ±1% of the reference value.

Required findings count

required_findings_count

Required finding count

Scores if within ±1% of the reference value.

Required information requests count

required_information_requests_count

Required information-request count

Scores if within ±1% of the reference value.

Required carried actions count

required_carried_actions_count

Required carried-action count

Scores if within ±1% of the reference value.

Difficulty

Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.

easy

Some inputs hidden

Visible provenance gaps and direct report identity conflicts

Hidden inputs

  • Packet variantpacket_variant
  • Previous area deltaprevious_area_delta_ha
  • Continuity margin targetcontinuity_margin_target_percent
  • Continuity deficitcontinuity_deficit_percent
  • Prior peak flow deltaprior_peak_flow_delta_m3_s
  • Prior hgl deltaprior_hgl_delta_m

Packet variant restricted to: clean, missing_manifest_catchment_revision, report_run_id_mismatch

medium

Some inputs hidden

Full provenance and propagation condition distribution

Hidden inputs

  • Packet variantpacket_variant
  • Previous area deltaprevious_area_delta_ha
  • Continuity margin targetcontinuity_margin_target_percent
  • Continuity deficitcontinuity_deficit_percent
  • Prior peak flow deltaprior_peak_flow_delta_m3_s
  • Prior hgl deltaprior_hgl_delta_m

hard

Some inputs hidden

Subtle revision, scenario, acceptance, and downstream propagation conflicts

Hidden inputs

  • Packet variantpacket_variant
  • Previous area deltaprevious_area_delta_ha
  • Continuity margin targetcontinuity_margin_target_percent
  • Continuity deficitcontinuity_deficit_percent
  • Prior peak flow deltaprior_peak_flow_delta_m3_s
  • Prior hgl deltaprior_hgl_delta_m

Packet variant restricted to: stale_catchment_revision, continuity_limit_exceeded, scenario_copy_forward, downstream_memo_stale_report, minor_open_comment_carried

Task bundle

The exact instruction and parameter contract used to generate this task, pinned to the published library source.

/workspace

  • instruction.md

Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.

1You are the independent reviewing engineer for a drainage model-run provenance package. The packet links governing catchment, rainfall, network, and configuration sources to a model manifest, run register, hydraulic report, and downstream design memo.2 3A task-owned synthetic source packet has been placed in `/workspace/sources/`. Treat it as the only source of truth. Your job is not to rebuild the drainage model. Your job is to decide which evidence can govern the design issue and to leave an auditable transition record.4 5## Review Workflow6 71. Inventory every registered source with document ID, revision, and status.82. Build a provenance ledger preserving the site, catchment, governing inputs, model manifest, reviewed run, report, design memo, and criteria source.93. Establish source authority before relying on dates or filenames. Trace the exact input revisions and scenario into the reviewed run.104. Recompute only the comparisons needed by the review matrix, using the source-owned assessment basis.115. Assign exactly one status to every review item: `pass`, `fail`, `not_applicable`, or `insufficient_data`.126. Do not invent missing values. Mark missing evidence as `insufficient_data` and request the exact missing field and source. A value explicitly marked pending or awaiting confirmation is missing evidence; it does not create unrelated identity failures.137. Convert every failure into one finding with a source pointer, affected object, consequence, and corrective action.148. Record the resulting state of the reviewed model run, its report, and the downstream design claim.159. Issue a readiness decision that reconciles with the matrix, transition decision, findings, information requests, and carried actions.16 17## Review Matrix18 19| Item | Review question |20|---|---|21| PRV-01 | Packet completeness: are all required input-basis, manifest, run-register, report, design-memo, and criteria records present with IDs and revisions? |22| PRV-02 | Source authority and identity: do site, catchment, governing source, run, report, memo, and criteria identities remain traceable without mixing objects? |23| PRV-03 | Input-revision provenance: does the reviewed run manifest identify the governing input revisions declared by the packet? |24| PRV-04 | Run/report integrity: does the report belong to the registered run and satisfy the packet's intrinsic report acceptance checks? The upstream input-governance state is recorded in the transition decision. |25| PRV-05 | Scenario propagation: is the governing design scenario preserved into the reviewed run? |26| PRV-06 | Downstream claim propagation: does the design memo preserve citation and value-propagation integrity for the reviewed run and report? Whether the cited evidence governs is recorded in the transition decision. |27| PRV-07 | Comment and action closure: is every critical comment closed, and is every carried action controlled by owner and action? |28| PRV-08 | Transition and readiness consistency: do the final states reconcile with the matrix and registers? |29| PRV-09 | Claim boundary: does the review avoid unsupported approval, compliance, source-hardening, executable-verifier, or benchmark-readiness claims? |30 31## Boundary Rules32 33- Use the matrix definitions and source packet to locate the most specific affected item. Do not infer a status from this instruction alone.34- If evidence needed for a comparison is absent, omit its `computed_evidence` key and raise an information request. Do not use `null`, zero, or placeholders for missing evidence.35- Do not double-count one source defect across unrelated matrix items. Consequences may propagate through the transition decision without becoming additional findings.36- Every finding, information request, and action must name one exact PRV item.37- Do not rename `computed_evidence` keys.38 39## Output40 41Write the complete review to `/workspace/output.md`. Explain the provenance decision briefly, then end with exactly one fenced JSON block:42 43```json44{45 "source_inventory": [{"doc_id": "...", "revision": "...", "status": "..."}],46 "provenance_ledger": {47 "site": "...",48 "catchment": "...",49 "catchment_basis": "...",50 "rainfall_basis": "...",51 "design_storm": "...",52 "network_model": "...",53 "model_config": "...",54 "run_manifest": "...",55 "reviewed_run": "...",56 "reviewed_report": "...",57 "design_memo": "...",58 "criteria_memo": "..."59 },60 "review_matrix": {61 "PRV-01": {"status": "pass|fail|not_applicable|insufficient_data", "evidence": "..."},62 "PRV-02": {"status": "...", "evidence": "..."},63 "PRV-03": {"status": "...", "evidence": "..."},64 "PRV-04": {"status": "...", "evidence": "..."},65 "PRV-05": {"status": "...", "evidence": "..."},66 "PRV-06": {"status": "...", "evidence": "..."},67 "PRV-07": {"status": "...", "evidence": "..."},68 "PRV-08": {"status": "...", "evidence": "..."},69 "PRV-09": {"status": "...", "evidence": "..."}70 },71 "computed_evidence": {72 "all_input_revisions_match_score": 0,73 "scenario_match_score": 0,74 "report_run_match_score": 0,75 "continuity_error_percent": 0,76 "continuity_margin_percent": 0,77 "report_peak_flow_m3_s": 0,78 "memo_peak_flow_m3_s": 0,79 "peak_flow_propagation_delta_m3_s": 0,80 "report_max_hgl_m_ahd": 0,81 "memo_max_hgl_m_ahd": 0,82 "hgl_propagation_delta_m": 083 },84 "transition_decision": {85 "model_run": "governing|non_governing|insufficient_data",86 "model_report": "governing|non_governing|insufficient_data",87 "design_claim": "supported|unsupported|insufficient_data"88 },89 "findings": [90 {"item": "PRV-0X", "severity": "critical|minor", "source_id": "...", "object_id": "...", "consequence": "...", "action": "..."}91 ],92 "information_requests": [93 {"item": "PRV-0X", "missing_field": "...", "source_id": "..."}94 ],95 "action_register": [96 {"action": "...", "owner": "...", "linked_item": "PRV-0X"}97 ],98 "readiness_decision": "ready_to_issue|ready_with_carried_actions|not_ready_to_issue",99 "claim_boundary_statement": "..."100}101```102 103Omit a `computed_evidence` key only when its required source field is missing, and raise the matching information request. Every `fail` needs a complete finding. Every `insufficient_data` needs an exact request. Every carried action needs an owner. The claim-boundary statement must say this is a task-owned synthetic source packet and does not claim authority approval, accepted project evidence, full standards compliance, source-pack hardening, executable-verifier readiness, or benchmark readiness.104

Scenario archetypes

Each generated task is drawn from one of these realistic scenario bands.

Site contexts ground each scenario in a real locale the model can use to infer hidden values.

Urban infill catchment

urban_infill_catchment

Urban infill catchment with a revised impervious-area basis

urban-infill-catchmentcity-centre-drainage-upgrade
Parameter ranges
current_catchment_area_ha
16 – 30
report_peak_flow_m3_s
1.2 – 3.2

Industrial precinct catchment

industrial_precinct_catchment

Industrial precinct catchment with staged network model revisions

industrial-precinct-catchmentbrownfield-drainage-upgrade
Parameter ranges
current_catchment_area_ha
28 – 48
report_peak_flow_m3_s
2.6 – 4.8

Example task

urban-infill-catchment-urban-infill-catchment-previewhard difficulty, some inputs hidden.

Urban infill catchment with a revised impervious-area basis. urban-infill-catchment. Required outputs: prv_01_status, prv_02_status, prv_03_status, prv_04_status, prv_05_status, prv_06_status

The model sees

Scenario context and visible inputs.

current_catchment_area_ha
16 to 30 ha
report_peak_flow_m3_s
1.2 to 3.2 m3/s
report_max_hgl_m_ahd
11.2 to 24.8 m AHD
maximum_continuity_error_percent
1.5 to 2.5 percent

Executable tool: drainage-model-run-provenance-issue-review-package_calc.py

The model must infer

Inputs withheld at this difficulty.

  • Prior peak flow delta

    prior_peak_flow_delta_m3_s

  • Prior hgl delta

    prior_hgl_delta_m

  • Packet variant

    packet_variant

  • Previous area delta

    previous_area_delta_ha

  • Continuity margin target

    continuity_margin_target_percent

  • Continuity deficit

    continuity_deficit_percent

The model must produce

The scored JSON answer schema.

{
  "prv_01_status": <number>,
  "prv_02_status": <number>,
  "prv_03_status": <number>,
  "prv_04_status": <number>,
  "prv_05_status": <number>,
  "prv_06_status": <number>,
  "prv_07_status": <number>,
  "prv_08_status": <number>,
  "prv_09_status": <number>,
  "all_input_revisions_match_score": <number>,
  "scenario_match_score": <number>,
  "report_run_match_score": <number>,
  "continuity_error_percent": <number>,
  "continuity_margin_percent": <number>,
  "report_peak_flow_m3_s": <number>,
  "memo_peak_flow_m3_s": <number>,
  "peak_flow_propagation_delta_m3_s": <number>,
  "report_max_hgl_m_ahd": <number>,
  "memo_max_hgl_m_ahd": <number>,
  "hgl_propagation_delta_m": <number>,
  "run_applicability_code": <number>,
  "report_applicability_code": <number>,
  "design_claim_support_code": <number>,
  "readiness_code": <number>,
  "required_findings_count": <number>,
  "required_information_requests_count": <number>,
  "required_carried_actions_count": <number>
}
  • prv_01_status · scored within ±1%
  • prv_02_status · scored within ±1%
  • prv_03_status · scored within ±1%
  • prv_04_status · scored within ±1%
  • prv_05_status · scored within ±1%
  • prv_06_status · scored within ±1%
  • prv_07_status · scored within ±1%
  • prv_08_status · scored within ±1%
  • prv_09_status · scored within ±1%
  • all_input_revisions_match_score · scored within ±1%
  • scenario_match_score · scored within ±1%
  • report_run_match_score · scored within ±1%
  • continuity_error_percent · scored within ±2%
  • continuity_margin_percent · scored within ±2%
  • report_peak_flow_m3_s · scored within ±2%
  • memo_peak_flow_m3_s · scored within ±2%
  • peak_flow_propagation_delta_m3_s · scored within ±2%
  • report_max_hgl_m_ahd · scored within ±2%
  • memo_max_hgl_m_ahd · scored within ±2%
  • hgl_propagation_delta_m · scored within ±2%
  • run_applicability_code · scored within ±1%
  • report_applicability_code · scored within ±1%
  • design_claim_support_code · scored within ±1%
  • readiness_code · scored within ±1%
  • required_findings_count · scored within ±1%
  • required_information_requests_count · scored within ±1%
  • required_carried_actions_count · scored within ±1%