LIVEdataset aec-bench@releasetasks 552models 18last submission · built
civilno-tool

Corridor Comment Response Issue Review Package

Review-native long-horizon task over a generated SSC-01 multimodal corridor comment-response packet. The agent receives a multi-file source packet in /workspace/sources/ (document register, comment register, marked-up plan, drainage recalculation, signal and pedestrian recalculation, VMS operations note, electrical feeder check, and criteria/comments), inventories it, preserves corridor identity, recomputes the package's own chainage, drainage, pedestrian, VMS, voltage-drop, and closeout evidence, assigns one status per review item, raises findings/information requests/actions, and issues a readiness decision. Packet variants inject missing revised chainage evidence, stale change-register revisions, chainage identity mismatch, scenario copy-forward, open comments, carried comments, or a genuine unsupported downstream repair. Scored by a stage-gated custom verifier, not per-key answer matching.

no-tool: The model must reason numerically unaided.

How this task is generated

One template produces many comparable benchmark tasks while keeping the scoring contract fixed.

  1. 01

    Template

    The reusable contract shown on this page.

  2. 02

    Scenario

    An archetype and site context are sampled.

  3. 03

    Difficulty tier

    Inputs may be hidden at harder tiers.

  4. 04

    Task prompt

    The model responds with the declared outputs.

Parameters

Inputs the model receives, and the outputs it is scored on.

Inputs

23 inputs

Always given

Included directly in every task prompt.

17
Show 17 inputs
  • Original chainage

    original_chainage_m

    C-008 original comment chainage on RD-SSC01-001

    1120 – 1620 m
  • Chainage delta

    chainage_delta_m

    Derived comment chainage movement from markup to revised package

    5 – 30 m
  • Revised hgl m

    revised_hgl_m

    DRN-08 revised HGL for the comment-response location

    18 – 48 m AHD
  • Minimum hgl clearance

    minimum_hgl_clearance_mm

    Criteria minimum road level to HGL clearance

    125 – 250 mm
  • Pedestrian startup time

    pedestrian_startup_time_s

    SG-08 pedestrian startup allowance

    3 – 4 s
  • Revised crossing width

    revised_crossing_width_m

    SG-08 revised pedestrian crossing width

    12 – 24 m
  • Pedestrian walk speed

    pedestrian_walk_speed_m_s

    SG-08 pedestrian walking speed

    1 – 1.3 m/s
  • Vms character height

    vms_character_height_in

    VMS-08 revised character height

    10.012.014.0
  • Approach speed kmh

    approach_speed_kmh

    CASE-08 approach speed used for VMS reading-time check

    40 – 60 km/h
  • Reading rate

    reading_rate_chars_s

    CASE-08 accepted VMS reading rate

    3.5 – 4.5 chars/s
  • Revised device load

    revised_device_load_w

    FEED-08 revised connected load after comment-response changes

    1200 – 3200 W
  • Feeder length

    feeder_length_km

    FEED-08 route length

    0.2 – 0.7 km
  • Conductor resistance

    conductor_resistance_ohm_km

    FEED-08 conductor resistance

    0.8 – 1.6 ohm/km
  • Feeder voltage

    feeder_voltage_v

    FEED-08 nominal feeder voltage

    230.0
  • Power factor

    power_factor

    FEED-08 load power factor

    0.85 – 0.95
  • Review comments total

    review_comments_total

    CMT-SSC01-008 total comments in the response batch

    5 – 8
  • Impacted calculation count

    impacted_calculation_count

    CMT-SSC01-008 count of calculations touched by the review change

    4 – 7

Hidden at higher difficulty

Visible in easier tasks and withheld in one or more harder tiers.

6
  • Hgl clearance margin

    hgl_clearance_margin_mm

    Hidden derivation margin between revised clearance and required HGL clearance

    Hidden at easy and medium and hard difficulty.

    60 – 260 mm
  • Ped clearance margin

    ped_clearance_margin_s

    Hidden derivation margin between available and required pedestrian clearance

    Hidden at easy and medium and hard difficulty.

    1 – 5 s
  • Vms message margin

    vms_message_margin_chars

    Hidden derivation margin between readable VMS capacity and message length

    Hidden at easy and medium and hard difficulty.

    2 – 10 chars
  • Voltage drop margin percent

    voltage_drop_margin_percent

    Hidden derivation margin between computed and allowable voltage drop

    Hidden at easy and medium and hard difficulty.

    0.8 – 4 %
  • Voltage drop deficit percent

    voltage_drop_deficit_percent

    Hidden derivation deficit for the unsupported-downstream-repair variant

    Hidden at easy and medium and hard difficulty.

    0.1 – 0.5 %
  • Packet variant

    packet_variant

    Hidden packet-defect variant controlling gold review statuses

    Hidden at easy and medium and hard difficulty.

    cleanmissing_revised_chainagestale_change_register_revisionchainage_identity_mismatchscenario_copy_forwardopen_critical_commentminor_open_comment_carriedunsupported_downstream_repair

Scored outputs

24 outputs

Rlr 01 status

rlr_01_status

Packet completeness status code (0 pass, 1 fail, 2 not applicable, 3 insufficient data)

Scores if within ±0% of the reference value.

Rlr 02 status

rlr_02_status

Object identity status code

Scores if within ±0% of the reference value.

Rlr 03 status

rlr_03_status

Review-response basis status code

Scores if within ±0% of the reference value.

Rlr 04 status

rlr_04_status

Review-response adequacy status code

Scores if within ±0% of the reference value.

Rlr 05 status

rlr_05_status

Scenario consequence status code

Scores if within ±0% of the reference value.

Rlr 06 status

rlr_06_status

Secondary-discipline resilience status code

Scores if within ±0% of the reference value.

Rlr 07 status

rlr_07_status

Comment and action closure status code

Scores if within ±0% of the reference value.

Rlr 08 status

rlr_08_status

Readiness decision consistency status code

Scores if within ±0% of the reference value.

Rlr 09 status

rlr_09_status

Claim boundary status code

Scores if within ±0% of the reference value.

Readiness code

readiness_code

Gold readiness decision (0 ready, 1 ready with carried actions, 2 not ready)

Scores if within ±0% of the reference value.

Required findings count

required_findings_count

Findings the review must raise

Scores if within ±0% of the reference value.

Required information requests count

required_information_requests_count

Information requests the review must raise

Scores if within ±0% of the reference value.

Required carried actions count

required_carried_actions_count

Carried actions the review must record

Scores if within ±0% of the reference value.

Changed chainage delta

changed_chainage_delta_m

Revised comment chainage minus original comment chainage

Scores if within ±2% of the reference value.

Hgl clearance

hgl_clearance_mm

Revised road level minus revised HGL

Scores if within ±2% of the reference value.

Hgl clearance margin

hgl_clearance_margin_mm

HGL clearance minus minimum required clearance

Scores if within ±2% of the reference value.

Ped clearance required s

ped_clearance_required_s

Required pedestrian clearance time

Scores if within ±2% of the reference value.

Ped clearance margin s

ped_clearance_margin_s

Available pedestrian clearance minus required clearance

Scores if within ±2% of the reference value.

Vms reading time s

vms_reading_time_s

Available VMS reading time

Scores if within ±2% of the reference value.

Vms message margin chars

vms_message_margin_chars

Readable character capacity minus revised message length

Scores if within ±2% of the reference value.

Feeder voltage drop

feeder_voltage_drop_percent

Voltage drop at revised device load

Scores if within ±2% of the reference value.

Voltage drop margin

voltage_drop_margin_percent

Allowable voltage drop minus computed voltage drop

Scores if within ±2% of the reference value.

Comment closeout

comment_closeout_percent

Closed comments divided by total comments

Scores if within ±2% of the reference value.

Impacted calculation count

impacted_calculation_count

Number of discipline calculations impacted by the review change

Scores if within ±2% of the reference value.

Difficulty

Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.

easy

Some inputs hidden

Obvious packet defects: clean, missing revised chainage, or a genuine downstream-repair failure

Hidden inputs

  • Packet variantpacket_variant
  • Hgl clearance marginhgl_clearance_margin_mm
  • Ped clearance margin sped_clearance_margin_s
  • Vms message margin charsvms_message_margin_chars
  • Voltage drop marginvoltage_drop_margin_percent
  • Voltage drop deficitvoltage_drop_deficit_percent

Packet variant restricted to: clean, missing_revised_chainage, unsupported_downstream_repair

medium

Some inputs hidden

Full packet-variant distribution

Hidden inputs

  • Packet variantpacket_variant
  • Hgl clearance marginhgl_clearance_margin_mm
  • Ped clearance margin sped_clearance_margin_s
  • Vms message margin charsvms_message_margin_chars
  • Voltage drop marginvoltage_drop_margin_percent
  • Voltage drop deficitvoltage_drop_deficit_percent

hard

Some inputs hidden

Subtle documentary defects: stale register, identity drift, copied scenario, and carried comments

Hidden inputs

  • Packet variantpacket_variant
  • Hgl clearance marginhgl_clearance_margin_mm
  • Ped clearance margin sped_clearance_margin_s
  • Vms message margin charsvms_message_margin_chars
  • Voltage drop marginvoltage_drop_margin_percent
  • Voltage drop deficitvoltage_drop_deficit_percent

Packet variant restricted to: stale_change_register_revision, chainage_identity_mismatch, scenario_copy_forward, minor_open_comment_carried

Task bundle

The exact instruction and parameter contract used to generate this task, pinned to the published library source.

/workspace

  • instruction.md

Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.

1You are the independent reviewing engineer for a multimodal corridor comment-response issue package covering one road corridor, one review comment, one drainage recalculation, one signal and pedestrian recalculation, one VMS operation note, and one ITS feeder check.2 3A source packet has been placed in `/workspace/sources/`. It contains a document register, comment register, marked-up plan and long section, drainage recalculation, signal and pedestrian recalculation, VMS operations note, electrical feeder check, and a criteria memo. The packet is a task-owned synthetic source pack; treat it as the only source of numeric truth for this review.4 5Your job is not to redesign the corridor. Your job is to decide whether the comment-response package is ready to issue, and to produce an auditable review record.6 7## Review Workflow8 91. Inventory the source packet before drawing any conclusion. Record every document ID, revision, and status.102. Build an identity ledger: corridor, comment, original and revised chainage, datum, scenario, drainage object, signal group, VMS device, and field feeder.113. Check for source conflicts, stale revisions, chainage drift, copied scenarios, missing evidence, and open critical comments before accepting any package claim.124. Recompute the package's own calculations only where they answer review items, using the assessment bases stated in the criteria memo. Do not import methods or values from outside the packet.135. Assign exactly one status to every review item: `pass`, `fail`, `not_applicable`, or `insufficient_data`.146. Do not invent missing values. Mark missing evidence as `insufficient_data` and request the exact missing field and source. A value that a source explicitly marks as pending or awaiting confirmation is missing evidence of this kind: it does not make otherwise-reconciling identifiers inconsistent and does not make evidence that is present untraceable; the check that cannot be completed without it takes `insufficient_data`.157. Convert every failure into a finding with a source pointer, affected object, consequence, and corrective action.168. Issue a readiness decision that reconciles with your matrix, findings, and action register.17 18## Review Discipline19 20Use the most specific review item for each defect based on the source packet and the review matrix definitions. Do not infer a status from this instruction alone.21 22- If a source value needed for a recomputation is absent, omit the dependent `computed_evidence` key and raise an information request for the missing field and source. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.23- Assign additional failures only when the source packet gives independent evidence for them; do not double-count one source issue across unrelated matrix rows.24- RLR-08 is reviewer self-consistency: it passes when your readiness decision reconciles with your matrix, findings, information requests, and action register.25- Every finding, information request, and action must name one exact RLR item.26 27## Review Matrix28 29Assess each item and give it exactly one status:30 31| Item | Review question |32|---|---|33| RLR-01 | Packet completeness: are all required source documents present with IDs and revisions? |34| RLR-02 | Object identity: do corridor, comment, chainage, datum, scenario, drainage, signal, VMS, and feeder identities stay consistent across documents? |35| RLR-03 | Review-response basis: are the change ledger and impacted calculation count traceable to current source revisions and recomputable evidence? |36| RLR-04 | Review-response adequacy: does the changed chainage/scenario propagate through HGL, pedestrian, VMS, voltage-drop, and comment-closeout checks? |37| RLR-05 | Scenario consequence: is the same corridor scenario used across disciplines rather than copied from another corridor? |38| RLR-06 | Secondary-discipline resilience: are VMS legibility and feeder voltage-drop checks source-backed and internally consistent? |39| RLR-07 | Comment and action closure: is every review comment closed, or carried with an owner and agreed action, or blocked by named missing data? |40| RLR-08 | Readiness consistency: does your final decision match your own matrix, findings, and action register? |41| RLR-09 | Claim boundary: does your review avoid unsupported approval, compliance, or acceptance claims? |42 43## Output44 45Write your complete review to `/workspace/output.md`. Explain your reasoning briefly in prose, then end with exactly one fenced JSON block:46 47```json48{49 "source_inventory": [{"doc_id": "...", "revision": "...", "status": "..."}],50 "identity_ledger": {51 "corridor": "...",52 "comment": "...",53 "original_chainage": "...",54 "revised_chainage": "...",55 "datum": "...",56 "scenario": "...",57 "drainage_object": "...",58 "signal_group": "...",59 "vms_device": "...",60 "field_feeder": "..."61 },62 "review_matrix": {63 "RLR-01": {"status": "pass|fail|not_applicable|insufficient_data", "evidence": "..."},64 "RLR-02": {"status": "...", "evidence": "..."},65 "RLR-03": {"status": "...", "evidence": "..."},66 "RLR-04": {"status": "...", "evidence": "..."},67 "RLR-05": {"status": "...", "evidence": "..."},68 "RLR-06": {"status": "...", "evidence": "..."},69 "RLR-07": {"status": "...", "evidence": "..."},70 "RLR-08": {"status": "...", "evidence": "..."},71 "RLR-09": {"status": "...", "evidence": "..."}72 },73 "computed_evidence": {74 "changed_chainage_delta_m": 0.0,75 "hgl_clearance_mm": 0.0,76 "hgl_clearance_margin_mm": 0.0,77 "ped_clearance_required_s": 0.0,78 "ped_clearance_margin_s": 0.0,79 "vms_reading_time_s": 0.0,80 "vms_message_margin_chars": 0.0,81 "feeder_voltage_drop_percent": 0.0,82 "voltage_drop_margin_percent": 0.0,83 "comment_closeout_percent": 0.0,84 "impacted_calculation_count": 0.085 },86 "findings": [87 {"item": "RLR-0X", "severity": "critical|minor", "source_id": "...", "object_id": "...", "consequence": "...", "action": "..."}88 ],89 "information_requests": [90 {"item": "RLR-0X", "missing_field": "...", "source_id": "..."}91 ],92 "action_register": [93 {"action": "...", "owner": "...", "linked_item": "RLR-0X"}94 ],95 "readiness_decision": "ready_to_issue|ready_with_carried_actions|not_ready_to_issue",96 "claim_boundary_statement": "..."97}98```99 100Rules for the structured block:101 102- Do not rename computed_evidence keys. Use the exact key names shown in the schema. Use `voltage_drop_margin_percent`, not `feeder_voltage_drop_margin_percent`.103- `computed_evidence` values must come from your own recomputation from packet source values. Omit a key only when its inputs are missing from the packet, then raise the matching information request instead. Do not include missing or unrecomputable evidence keys with `null`, `0`, or placeholder values.104- Every `fail` needs at least one finding with a non-empty `source_id`, `object_id`, `consequence`, and `action`. Each `findings[].item` value must be a single RLR item such as `RLR-05`; do not combine multiple IDs in one item field.105- Every `insufficient_data` needs an information request naming the exact missing field and its source document. Each `information_requests[].item` value must be a single RLR item such as `RLR-04`.106- Every `not_applicable` needs a scope reason in its matrix `evidence`.107- Carried actions must appear in `action_register` with an owner. Each `action_register[].linked_item` value must be a single RLR item.108- `readiness_decision` must reconcile with your matrix: unresolved failures or missing critical evidence mean the package is not ready.109- `claim_boundary_statement` must state that this review covers a task-owned synthetic source packet and does not claim authority approval, accepted project evidence, full standards compliance, source-pack hardening, executable-verifier readiness, or benchmark readiness.110

Scenario archetypes

Each generated task is drawn from one of these realistic scenario bands.

Site contexts ground each scenario in a real locale the model can use to infer hidden values.

Urban arterial response

urban_arterial_response

Urban arterial comment-response package with wider crossing and larger ITS load

urban-arterial-comment-responsecbd-fringe-corridor-review
Parameter ranges
revised_crossing_width_m
16 – 24
revised_device_load_w
1900 – 3200
approach_speed_kmh
45 – 60

Suburban collector response

suburban_collector_response

Suburban collector comment-response package with moderate crossing width and shorter feeder

suburban-collector-comment-responseschool-access-corridor-review
Parameter ranges
revised_crossing_width_m
12 – 19
revised_device_load_w
1200 – 2400
feeder_length_km
0.2 – 0.45

Example task

urban-arterial-comment-response-urban-arterial-response-previewhard difficulty, some inputs hidden.

Urban arterial comment-response package with wider crossing and larger ITS load. urban-arterial-comment-response. Required outputs: rlr_01_status, rlr_02_status, rlr_03_status, rlr_04_status, rlr_05_status, rlr_06_status

The model sees

Scenario context and visible inputs.

original_chainage_m
1120 to 1620 m
chainage_delta_m
5 to 30 m
revised_hgl_m
18 to 48 m AHD
minimum_hgl_clearance_mm
125 to 250 mm
pedestrian_startup_time_s
3 to 4 s
revised_crossing_width_m
16 to 24 m
pedestrian_walk_speed_m_s
1 to 1.3 m/s
vms_character_height_in
10.0 in
approach_speed_kmh
45 to 60 km/h
reading_rate_chars_s
3.5 to 4.5 chars/s
revised_device_load_w
1900 to 3200 W
feeder_length_km
0.2 to 0.7 km
conductor_resistance_ohm_km
0.8 to 1.6 ohm/km
feeder_voltage_v
230.0 V
power_factor
0.85 to 0.95
review_comments_total
5 to 8
impacted_calculation_count
4 to 7

Executable tool: corridor-comment-response-issue-review-package_calc.py

The model must infer

Inputs withheld at this difficulty.

  • Ped clearance margin s

    ped_clearance_margin_s

  • Packet variant

    packet_variant

  • Voltage drop deficit

    voltage_drop_deficit_percent

  • Vms message margin chars

    vms_message_margin_chars

  • Voltage drop margin

    voltage_drop_margin_percent

  • Hgl clearance margin

    hgl_clearance_margin_mm

The model must produce

The scored JSON answer schema.

{
  "rlr_01_status": <number>,
  "rlr_02_status": <number>,
  "rlr_03_status": <number>,
  "rlr_04_status": <number>,
  "rlr_05_status": <number>,
  "rlr_06_status": <number>,
  "rlr_07_status": <number>,
  "rlr_08_status": <number>,
  "rlr_09_status": <number>,
  "readiness_code": <number>,
  "required_findings_count": <number>,
  "required_information_requests_count": <number>,
  "required_carried_actions_count": <number>,
  "changed_chainage_delta_m": <number>,
  "hgl_clearance_mm": <number>,
  "hgl_clearance_margin_mm": <number>,
  "ped_clearance_required_s": <number>,
  "ped_clearance_margin_s": <number>,
  "vms_reading_time_s": <number>,
  "vms_message_margin_chars": <number>,
  "feeder_voltage_drop_percent": <number>,
  "voltage_drop_margin_percent": <number>,
  "comment_closeout_percent": <number>,
  "impacted_calculation_count": <number>
}
  • rlr_01_status · scored within ±0%
  • rlr_02_status · scored within ±0%
  • rlr_03_status · scored within ±0%
  • rlr_04_status · scored within ±0%
  • rlr_05_status · scored within ±0%
  • rlr_06_status · scored within ±0%
  • rlr_07_status · scored within ±0%
  • rlr_08_status · scored within ±0%
  • rlr_09_status · scored within ±0%
  • readiness_code · scored within ±0%
  • required_findings_count · scored within ±0%
  • required_information_requests_count · scored within ±0%
  • required_carried_actions_count · scored within ±0%
  • changed_chainage_delta_m · scored within ±2%
  • hgl_clearance_mm · scored within ±2%
  • hgl_clearance_margin_mm · scored within ±2%
  • ped_clearance_required_s · scored within ±2%
  • ped_clearance_margin_s · scored within ±2%
  • vms_reading_time_s · scored within ±2%
  • vms_message_margin_chars · scored within ±2%
  • feeder_voltage_drop_percent · scored within ±2%
  • voltage_drop_margin_percent · scored within ±2%
  • comment_closeout_percent · scored within ±2%
  • impacted_calculation_count · scored within ±2%