LIVEdataset aec-bench@releasetasks 552models 18last submission · built
civilno-tool

Coastal Flood Equipment Elevation Issue Review Package

Review-first source-packet task for a coastal pump-station equipment elevation package: inventory source files, preserve datum and planning-horizon identity, recompute design flood level and equipment margins, and issue an auditable readiness decision.

no-tool: The model must reason numerically unaided.

How this task is generated

One template produces many comparable benchmark tasks while keeping the scoring contract fixed.

  1. 01

    Template

    The reusable contract shown on this page.

  2. 02

    Scenario

    An archetype and site context are sampled.

  3. 03

    Difficulty tier

    Inputs may be hidden at harder tiers.

  4. 04

    Task prompt

    The model responds with the declared outputs.

Parameters

Inputs the model receives, and the outputs it is scored on.

Inputs

14 inputs

Always given

Included directly in every task prompt.

8
Show 8 inputs
  • Present mean sea level

    present_mean_sea_level_m_ahd

    Present mean sea level in AHD

    0.1 – 0.35 m AHD
  • Tide amplitude

    tide_amplitude_m

    Design tide amplitude

    0.75 – 1.05 m
  • Storm surge

    storm_surge_m

    Storm surge allowance

    0.9 – 1.3 m
  • Slr allowance

    slr_allowance_m

    Sea-level-rise allowance for the selected planning horizon

    0.4 – 0.7 m
  • Wave runup

    wave_runup_m

    Wave/runup allowance

    0.3 – 0.55 m
  • Required freeboard

    required_freeboard_m

    Required equipment freeboard above design flood level

    0.3 – 0.45 m
  • Required pump duty

    required_pump_duty_m3_s

    Required pump duty during selected coastal event

    0.3 – 0.48 m3/s
  • Cd to ahd offset

    cd_to_ahd_offset_m

    Chart datum to AHD offset owned by the criteria memo

    -0.92 – -0.72 m

Hidden at higher difficulty

Visible in easier tasks and withheld in one or more harder tiers.

6
  • Switchboard margin target

    switchboard_margin_target_m

    Hidden clean switchboard freeboard margin used to derive elevation

    Hidden at easy and medium and hard difficulty.

    0.18 – 0.42 m
  • Generator margin target

    generator_margin_target_m

    Hidden generator freeboard margin used to derive elevation

    Hidden at easy and medium and hard difficulty.

    0.25 – 0.55 m
  • Switchboard deficit

    switchboard_deficit_m

    Hidden deficient-case switchboard freeboard deficit

    Hidden at easy and medium and hard difficulty.

    0.08 – 0.25 m
  • Outfall tailwater margin

    outfall_tailwater_margin_m

    Hidden outfall allowable tailwater margin

    Hidden at easy and medium and hard difficulty.

    0.2 – 0.55 m
  • Pump duty margin target

    pump_duty_margin_target_m3_s

    Hidden selected pump duty margin

    Hidden at easy and medium and hard difficulty.

    0.04 – 0.12 m3/s
  • Packet variant

    packet_variant

    Hidden source-packet issue variant

    Hidden at easy and medium and hard difficulty.

    cleanmissing_switchboard_survey_levelstale_water_level_basis_revisionasset_survey_chart_datum_labelled_ahdscenario_copy_forwardopen_critical_commentminor_open_comment_carriedswitchboard_below_design_level

Scored outputs

20 outputs

Rlr 01 status

rlr_01_status

RLR-01 status code

Scores if within ±1% of the reference value.

Rlr 02 status

rlr_02_status

RLR-02 status code

Scores if within ±1% of the reference value.

Rlr 03 status

rlr_03_status

RLR-03 status code

Scores if within ±1% of the reference value.

Rlr 04 status

rlr_04_status

RLR-04 status code

Scores if within ±1% of the reference value.

Rlr 05 status

rlr_05_status

RLR-05 status code

Scores if within ±1% of the reference value.

Rlr 06 status

rlr_06_status

RLR-06 status code

Scores if within ±1% of the reference value.

Rlr 07 status

rlr_07_status

RLR-07 status code

Scores if within ±1% of the reference value.

Rlr 08 status

rlr_08_status

RLR-08 status code

Scores if within ±1% of the reference value.

Rlr 09 status

rlr_09_status

RLR-09 status code

Scores if within ±1% of the reference value.

Design flood level m ahd

design_flood_level_m_ahd

Design flood level in AHD

Scores if within ±3% of the reference value.

Wave runup

wave_runup_m

Wave/runup allowance evidence

Scores if within ±3% of the reference value.

Switchboard freeboard

switchboard_freeboard_m

Switchboard elevation minus design flood level

Scores if within ±3% of the reference value.

Switchboard freeboard margin

switchboard_freeboard_margin_m

Switchboard freeboard minus required freeboard

Scores if within ±3% of the reference value.

Generator freeboard margin

generator_freeboard_margin_m

Generator freeboard minus required freeboard

Scores if within ±3% of the reference value.

Outfall submergence margin

outfall_submergence_margin_m

Outfall allowable tailwater minus design tailwater

Scores if within ±3% of the reference value.

Pump duty margin

pump_duty_margin

Selected pump duty minus required pump duty

Scores if within ±3% of the reference value.

Readiness code

readiness_code

Readiness decision code

Scores if within ±1% of the reference value.

Required findings count

required_findings_count

Required finding count

Scores if within ±1% of the reference value.

Required information requests count

required_information_requests_count

Required information request count

Scores if within ±1% of the reference value.

Required carried actions count

required_carried_actions_count

Required carried action count

Scores if within ±1% of the reference value.

Difficulty

Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.

easy

Some inputs hidden

Obvious coastal equipment-elevation defects

Hidden inputs

  • Packet variantpacket_variant
  • Switchboard margin targetswitchboard_margin_target_m
  • Generator margin targetgenerator_margin_target_m
  • Switchboard deficitswitchboard_deficit_m
  • Outfall tailwater marginoutfall_tailwater_margin_m
  • Pump duty margin targetpump_duty_margin_target_m3_s

Packet variant restricted to: clean, missing_switchboard_survey_level, switchboard_below_design_level

medium

Some inputs hidden

Full packet-variant distribution

Hidden inputs

  • Packet variantpacket_variant
  • Switchboard margin targetswitchboard_margin_target_m
  • Generator margin targetgenerator_margin_target_m
  • Switchboard deficitswitchboard_deficit_m
  • Outfall tailwater marginoutfall_tailwater_margin_m
  • Pump duty margin targetpump_duty_margin_target_m3_s

hard

Some inputs hidden

Subtle datum, revision, planning-horizon, and comment-closure defects

Hidden inputs

  • Packet variantpacket_variant
  • Switchboard margin targetswitchboard_margin_target_m
  • Generator margin targetgenerator_margin_target_m
  • Switchboard deficitswitchboard_deficit_m
  • Outfall tailwater marginoutfall_tailwater_margin_m
  • Pump duty margin targetpump_duty_margin_target_m3_s

Packet variant restricted to: stale_water_level_basis_revision, asset_survey_chart_datum_labelled_ahd, scenario_copy_forward, minor_open_comment_carried

Task bundle

The exact instruction and parameter contract used to generate this task, pinned to the published library source.

/workspace

  • instruction.md

Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.

1You are the independent reviewing engineer for a coastal flood equipment elevation issue package covering one pump station/outfall, tide and water-level basis, SLR planning horizon, wave/runup basis, asset survey, pump/outfall schedule, and criteria/comments memo.2 3A source packet has been placed in `/workspace/sources/`. It contains a document register, tide and water-level basis, SLR planning-horizon table, wave/runup basis, asset survey, pump/outfall schedule, and criteria memo with review comments. The packet is a task-owned synthetic source pack; treat it as the only source of numeric truth for this review.4 5Your job is not to redesign the coastal protection, pump station, or electrical equipment. Your job is to decide whether the package is ready to issue, and to produce an auditable review record.6 7## Review Workflow8 91. Inventory the source packet before drawing any conclusion. Record every document ID, revision, and status.102. Build an identity ledger: site, datum, water-level basis, SLR scenario, runup case, switchboard, generator, outfall, pump duty, and criteria memo.113. Check for source conflicts, stale revisions, copied scenarios, missing evidence, datum mismatches, and open critical comments before accepting any package claim.124. Recompute the package's own calculations only where they answer review items, using the assessment bases stated in the criteria memo. Do not import methods or values from outside the packet.135. Assign exactly one status to every review item: `pass`, `fail`, `not_applicable`, or `insufficient_data`.146. Do not invent missing values. Mark missing evidence as `insufficient_data` and request the exact missing field and source. A value that a source explicitly marks as pending or awaiting confirmation is missing evidence of this kind: it does not make otherwise-reconciling identifiers inconsistent and does not make evidence that is present untraceable; the check that cannot be completed without it takes `insufficient_data`.157. Convert every failure into a finding with a source pointer, affected object, consequence, and corrective action.168. Issue a readiness decision that reconciles with your matrix, findings, information requests, and action register.17 18## Review Matrix19 20Assess each item and give it exactly one status:21 22| Item | Review question |23|---|---|24| RLR-01 | Packet completeness: are all required tide/water-level, SLR, runup, survey, pump/outfall, and criteria files present with IDs and revisions? |25| RLR-02 | Object identity and datum consistency: do site, datum, switchboard, generator, outfall, pump, and criteria references stay consistent without mixing declared datum frames? |26| RLR-03 | Coastal boundary basis: are tide, SLR, surge, runup, datum offset, and design flood level traceable, current, and recomputable? |27| RLR-04 | Equipment elevation adequacy: do switchboard and generator freeboard margins clear the source-owned design flood level and required freeboard? |28| RLR-05 | Scenario consequence: is the same planning horizon, asset class, and coastal event used across water-level, SLR, runup, survey, and pump/outfall sources? |29| RLR-06 | Outfall and pump resilience: are outfall submergence and pump duty margins source-backed and internally consistent with the selected coastal event? |30| RLR-07 | Comment and action closure: is every critical comment closed, and are carried minor comments owner/action controlled? |31| RLR-08 | Readiness consistency: does your final decision match your own matrix, findings, information requests, and action register? |32| RLR-09 | Claim boundary: does your review avoid unsupported approval, compliance, source-hardening, executable-verifier, or benchmark-readiness claims? |33 34## Boundary Rules35 36- Use the review matrix definitions to decide the most specific affected item from the source packet. Do not infer a status from this instruction alone.37- If a source value needed for a recomputation is absent, omit the dependent `computed_evidence` key and raise an information request for the missing field and source. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.38- Assign additional failures only when the source packet gives independent evidence for them; do not double-count one source issue across unrelated matrix rows.39- RLR-08 passes when the readiness decision reconciles with the review matrix, findings, information requests, and action register.40- Every finding, information request, and action must name one exact RLR item. Use a single RLR item per register row; do not write combined items such as `RLR-04/RLR-06`.41- Do not rename computed_evidence keys.42 43## Output44 45Write your complete review to `/workspace/output.md`. Explain your reasoning briefly in prose, then end with exactly one fenced JSON block:46 47```json48{49 "source_inventory": [{"doc_id": "...", "revision": "...", "status": "..."}],50 "identity_ledger": {51 "site": "...",52 "datum": "...",53 "water_level_basis": "...",54 "slr_scenario": "...",55 "runup_case": "...",56 "switchboard": "...",57 "generator": "...",58 "outfall": "...",59 "pump_duty": "...",60 "criteria_memo": "..."61 },62 "review_matrix": {63 "RLR-01": {"status": "pass|fail|not_applicable|insufficient_data", "evidence": "..."},64 "RLR-02": {"status": "...", "evidence": "..."},65 "RLR-03": {"status": "...", "evidence": "..."},66 "RLR-04": {"status": "...", "evidence": "..."},67 "RLR-05": {"status": "...", "evidence": "..."},68 "RLR-06": {"status": "...", "evidence": "..."},69 "RLR-07": {"status": "...", "evidence": "..."},70 "RLR-08": {"status": "...", "evidence": "..."},71 "RLR-09": {"status": "...", "evidence": "..."}72 },73 "computed_evidence": {74 "design_flood_level_m_ahd": 0.0,75 "wave_runup_m": 0.0,76 "switchboard_freeboard_m": 0.0,77 "switchboard_freeboard_margin_m": 0.0,78 "generator_freeboard_margin_m": 0.0,79 "outfall_submergence_margin_m": 0.0,80 "pump_duty_margin": 0.081 },82 "findings": [83 {"item": "RLR-0X", "severity": "critical|minor", "source_id": "...", "object_id": "...", "consequence": "...", "action": "..."}84 ],85 "information_requests": [86 {"item": "RLR-0X", "missing_field": "...", "source_id": "..."}87 ],88 "action_register": [89 {"action": "...", "owner": "...", "linked_item": "RLR-0X"}90 ],91 "readiness_decision": "ready_to_issue|ready_with_carried_actions|not_ready_to_issue",92 "claim_boundary_statement": "..."93}94```95 96Rules for the structured block:97 98- `computed_evidence` values must come from your own recomputation from packet source values. Omit a key only when its inputs are missing from the packet, then raise the matching information request instead. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.99- Every `fail` needs at least one finding with a non-empty `source_id`, `object_id`, `consequence`, and `action`.100- Every `insufficient_data` needs an information request naming the exact missing field and its source document.101- Every `not_applicable` needs a scope reason in its matrix `evidence`.102- Carried actions must appear in `action_register` with an owner.103- `readiness_decision` must reconcile with your matrix: unresolved failures or missing critical evidence mean the package is not ready.104- `claim_boundary_statement` must state that this review covers a task-owned synthetic source packet and does not claim authority approval, accepted project evidence, full standards compliance, source-pack hardening, executable-verifier readiness, or benchmark readiness.105

Scenario archetypes

Each generated task is drawn from one of these realistic scenario bands.

Site contexts ground each scenario in a real locale the model can use to infer hidden values.

Estuary pump station

estuary_pump_station

Estuary pump station with AHD-controlled equipment elevations

estuary-pump-stationcoastal-utility-yard
Parameter ranges
present_mean_sea_level_m_ahd
0.1 – 0.25
tide_amplitude_m
0.75 – 0.95
storm_surge_m
0.9 – 1.15

Harbour outfall station

harbour_outfall_station

Harbour outfall pump station with higher tide and runup allowances

harbour-outfall-stationmarine-edge-pump-station
Parameter ranges
present_mean_sea_level_m_ahd
0.2 – 0.35
tide_amplitude_m
0.88 – 1.05
storm_surge_m
1.05 – 1.3

Example task

estuary-pump-station-estuary-pump-station-previewhard difficulty, some inputs hidden.

Estuary pump station with AHD-controlled equipment elevations. estuary-pump-station. Required outputs: rlr_01_status, rlr_02_status, rlr_03_status, rlr_04_status, rlr_05_status, rlr_06_status

The model sees

Scenario context and visible inputs.

present_mean_sea_level_m_ahd
0.1 to 0.25 m AHD
tide_amplitude_m
0.75 to 0.95 m
storm_surge_m
0.9 to 1.15 m
slr_allowance_m
0.4 to 0.7 m
wave_runup_m
0.3 to 0.55 m
required_freeboard_m
0.3 to 0.45 m
required_pump_duty_m3_s
0.3 to 0.48 m3/s
cd_to_ahd_offset_m
-0.92 to -0.72 m

Executable tool: coastal-flood-equipment-elevation-issue-review-package_calc.py

The model must infer

Inputs withheld at this difficulty.

  • Generator margin target

    generator_margin_target_m

  • Packet variant

    packet_variant

  • Pump duty margin target

    pump_duty_margin_target_m3_s

  • Outfall tailwater margin

    outfall_tailwater_margin_m

  • Switchboard margin target

    switchboard_margin_target_m

  • Switchboard deficit

    switchboard_deficit_m

The model must produce

The scored JSON answer schema.

{
  "rlr_01_status": <number>,
  "rlr_02_status": <number>,
  "rlr_03_status": <number>,
  "rlr_04_status": <number>,
  "rlr_05_status": <number>,
  "rlr_06_status": <number>,
  "rlr_07_status": <number>,
  "rlr_08_status": <number>,
  "rlr_09_status": <number>,
  "design_flood_level_m_ahd": <number>,
  "wave_runup_m": <number>,
  "switchboard_freeboard_m": <number>,
  "switchboard_freeboard_margin_m": <number>,
  "generator_freeboard_margin_m": <number>,
  "outfall_submergence_margin_m": <number>,
  "pump_duty_margin": <number>,
  "readiness_code": <number>,
  "required_findings_count": <number>,
  "required_information_requests_count": <number>,
  "required_carried_actions_count": <number>
}
  • rlr_01_status · scored within ±1%
  • rlr_02_status · scored within ±1%
  • rlr_03_status · scored within ±1%
  • rlr_04_status · scored within ±1%
  • rlr_05_status · scored within ±1%
  • rlr_06_status · scored within ±1%
  • rlr_07_status · scored within ±1%
  • rlr_08_status · scored within ±1%
  • rlr_09_status · scored within ±1%
  • design_flood_level_m_ahd · scored within ±3%
  • wave_runup_m · scored within ±3%
  • switchboard_freeboard_m · scored within ±3%
  • switchboard_freeboard_margin_m · scored within ±3%
  • generator_freeboard_margin_m · scored within ±3%
  • outfall_submergence_margin_m · scored within ±3%
  • pump_duty_margin · scored within ±3%
  • readiness_code · scored within ±1%
  • required_findings_count · scored within ±1%
  • required_information_requests_count · scored within ±1%
  • required_carried_actions_count · scored within ±1%