Rlr 01 status
rlr_01_status
Packet completeness status code (0 pass, 1 fail, 2 not applicable, 3 insufficient data)
Scores if within ±0% of the reference value.
Review-native long-horizon task over a generated SSC-01 road low-point issue packet. The agent receives a multi-file source packet in /workspace/sources/ (document register, road geometry, drainage package, field equipment, power/comms, traffic operations, criteria and comments), inventories it, preserves corridor identity, checks the package's own calculations as review evidence, assigns one status per review item, raises findings/information requests/actions, and issues a readiness decision. Packet variants inject missing evidence, stale revisions, datum mismatches, scenario copy-forward, open comments, or genuine criterion failures. Scored by a stage-gated custom verifier, not per-key answer matching.
no-tool: The model must reason numerically unaided.
Standards
One template produces many comparable benchmark tasks while keeping the scoring contract fixed.
01
The reusable contract shown on this page.
02
An archetype and site context are sampled.
03
Inputs may be hidden at harder tiers.
04
The model responds with the declared outputs.
Inputs the model receives, and the outputs it is scored on.
36 inputs
Included directly in every task prompt.
Runoff coefficient
runoff_coefficient
DRN-CAT-01 runoff coefficient for the contributing low-point catchment
Rainfall intensity
rainfall_intensity_mm_h
STORM-01 rainfall intensity for the low-point design burst
Catchment area
catchment_area_ha
DRN-CAT-01 contributing catchment area draining to LP-01
Upstream bypass flow
upstream_bypass_flow_m3_s
DRN-BYP-01 upstream inlet bypass flow reaching LP-01
Cross slope pct
cross_slope_pct
RD-SSC01-001 pavement crossfall at LP-01
Longitudinal slope pct
longitudinal_slope_pct
RD-SSC01-001 approach longitudinal slope into LP-01
Gutter mannings n
gutter_mannings_n
DRN-GUT-01 Manning roughness for the pavement gutter
Road low point level
road_low_point_level_m
LP-01 pavement level at the sag point (AHD)
Inlet efficiency
inlet_efficiency
DRN-PIT-01 low-point inlet interception efficiency
Inlet capture capacity
inlet_capture_capacity_m3_s
DRN-PIT-01 maximum inlet capture capacity
Pipe diameter
pipe_diameter_mm
DRN-PIPE-01 storm-drain diameter leaving DRN-PIT-01
Pipe length
pipe_length_m
DRN-PIPE-01 reach length to the downstream tailwater node
Pipe mannings n
pipe_mannings_n
DRN-PIPE-01 full-pipe Manning roughness
Pit loss coefficient
pit_loss_coefficient
HGL-01 pit and junction loss coefficient at DRN-PIT-01
Minimum cabinet freeboard
minimum_cabinet_freeboard_m
CRIT-SSC01-001 minimum freeboard from the controlling water level to the CAB-01 pad
Road design speed kmh
road_design_speed_kmh
RD-SSC01-001 design speed at the low point
Vms character height
vms_character_height_in
VMS-01 character height
Reading rate
reading_rate_chars_s
TOPS-SSC01-CASE-01 driver reading rate for VMS legibility
Camera count
camera_count
ITS-NET-01 CCTV cameras served from CAB-01
Camera data rate
camera_data_rate_mbps
ITS-NET-01 data rate per CCTV camera
Vms data rate
vms_data_rate_mbps
ITS-NET-01 VMS-01 data rate
Controller data rate
controller_data_rate_mbps
ITS-NET-01 controller data rate
Sensor data rate
sensor_data_rate_mbps
ITS-NET-01 water-level sensor data rate
Network overhead pct
network_overhead_pct
ITS-NET-01 protocol overhead allowance
Future capacity buffer pct
future_capacity_buffer_pct
ITS-NET-01 future capacity buffer
Required autonomy
required_autonomy_h
CRIT-SSC01-001 required battery autonomy for CAB-01 critical load
Battery efficiency
battery_efficiency
BATT-01 usable-energy efficiency
Critical load
critical_load_w
CAB-01 critical load during the event
Visible in easier tasks and withheld in one or more harder tiers.
Allowable spread margin
allowable_spread_margin_m
Derivation margin between computed spread and the CRIT-SSC01-001 allowable spread
Hidden at easy and medium and hard difficulty.
Hgl clearance target
hgl_clearance_target_m
Derivation clearance between the upstream HGL and the LP-01 pavement level
Hidden at easy and medium and hard difficulty.
Cabinet freeboard margin
cabinet_freeboard_margin_m
Derivation margin above minimum freeboard used to set the CAB-01 pad level
Hidden at easy and medium and hard difficulty.
Cabinet freeboard deficit
cabinet_freeboard_deficit_m
Derivation deficit below minimum freeboard for the freeboard-deficient variant
Hidden at easy and medium and hard difficulty.
Vms message margin chars target
vms_message_margin_chars_target
Derivation margin between readable characters and the selected message length
Hidden at easy and medium and hard difficulty.
Network headroom margin
network_headroom_margin_mbps
Derivation margin between required network load and the provisioned uplink
Hidden at easy and medium and hard difficulty.
Battery runtime margin
battery_runtime_margin_h
Derivation margin between required autonomy and provisioned battery runtime
Hidden at easy and medium and hard difficulty.
Packet variant
packet_variant
Hidden packet-defect variant controlling gold review statuses
Hidden at easy and medium and hard difficulty.
22 outputs
rlr_01_status
Packet completeness status code (0 pass, 1 fail, 2 not applicable, 3 insufficient data)
Scores if within ±0% of the reference value.
rlr_02_status
Object identity status code
Scores if within ±0% of the reference value.
rlr_03_status
Drainage basis status code
Scores if within ±0% of the reference value.
rlr_04_status
Equipment exposure status code
Scores if within ±0% of the reference value.
rlr_05_status
Traffic operation consequence status code
Scores if within ±0% of the reference value.
rlr_06_status
Power/comms resilience status code
Scores if within ±0% of the reference value.
rlr_07_status
Comment and action closure status code
Scores if within ±0% of the reference value.
rlr_08_status
Readiness decision consistency status code
Scores if within ±0% of the reference value.
rlr_09_status
Claim boundary status code
Scores if within ±0% of the reference value.
readiness_code
Gold readiness decision (0 ready, 1 ready with carried actions, 2 not ready)
Scores if within ±0% of the reference value.
required_findings_count
Findings the review must raise
Scores if within ±0% of the reference value.
required_information_requests_count
Information requests the review must raise
Scores if within ±0% of the reference value.
required_carried_actions_count
Carried actions the review must record
Scores if within ±0% of the reference value.
peak_runoff_m3_s
Rational-method peak runoff evidence
Scores if within ±2% of the reference value.
gutter_approach_flow_m3_s
Gutter approach flow evidence
Scores if within ±2% of the reference value.
spread_width_m
Triangular gutter spread evidence
Scores if within ±2% of the reference value.
allowable_spread_m
Criteria allowable spread (source-owned)
Scores if within ±0% of the reference value.
controlling_water_level_m
Controlling water level evidence
Scores if within ±2% of the reference value.
cabinet_freeboard_m
Cabinet freeboard evidence (absent when the pad level is missing)
Scores if within ±2% of the reference value.
vms_message_margin_chars
VMS message margin evidence
Scores if within ±2% of the reference value.
battery_runtime_h
Battery runtime evidence
Scores if within ±2% of the reference value.
network_headroom_mbps
Network headroom evidence
Scores if within ±2% of the reference value.
Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.
Some inputs hidden
Obvious packet defects: clean, missing evidence, or a genuine criterion failure
Hidden inputs
Packet variant restricted to: clean, missing_cabinet_level, freeboard_deficient
Some inputs hidden
Full packet-variant distribution
Hidden inputs
Some inputs hidden
Subtle documentary defects: stale revisions, datum drift, scenario copy-forward, carried comments
Hidden inputs
Packet variant restricted to: stale_hgl_revision, chainage_datum_mismatch, scenario_copy_forward, minor_open_comment_carried
The exact instruction and parameter contract used to generate this task, pinned to the published library source.
/workspace
Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.
1You are the independent reviewing engineer for a road-corridor issue package covering a sag low point near roadside field equipment.2 3A source packet has been placed in `/workspace/sources/`. It contains a document register, road geometry, a drainage design package, a field equipment layout, a power and network schedule, a traffic operations case, and a criteria memo with review comments. The packet is a task-owned synthetic source pack; treat it as the only source of numeric truth for this review.4 5Your job is not to redesign the package. Your job is to decide whether it is ready to issue, and to produce an auditable review record.6 7## Review Workflow8 91. Inventory the source packet before drawing any conclusion. Record every document ID, revision, and status.102. Build a corridor identity ledger: road segment, chainage frame, datum, low point, cabinet, VMS, storm case, network case, and battery.113. Check for source conflicts, stale revisions, contradictory datums, and missing evidence first.124. Recompute the package's own calculations only where they answer review items, using the assessment bases stated in the criteria memo. Do not import methods or values from outside the packet.135. Assign exactly one status to every review item: `pass`, `fail`, `not_applicable`, or `insufficient_data`.146. Do not invent missing values. Mark missing evidence as `insufficient_data` and request the exact missing field and source. A value that a source explicitly marks as pending or awaiting confirmation is missing evidence of this kind: it does not make otherwise-reconciling identifiers inconsistent and does not make evidence that is present untraceable; the check that cannot be completed without it takes `insufficient_data`.157. Convert every failure into a finding with a source pointer, affected object, consequence, and corrective action.168. Issue a readiness decision that reconciles with your matrix, findings, and action register.17 18## Review Matrix19 20Assess each item and give it exactly one status:21 22| Item | Review question |23|---|---|24| RLR-01 | Packet completeness: are all required source documents present with IDs and revisions? |25| RLR-02 | Object identity: do chainage, datum, low point, cabinet, storm case, and scenario stay consistent across documents? |26| RLR-03 | Drainage basis: are the package's runoff, spread, and HGL results traceable to the stated storm case and current tailwater basis, and do they reconcile when recomputed? |27| RLR-04 | Equipment exposure: is the CAB-01 pad level adequate against the controlling water level and the required freeboard? |28| RLR-05 | Traffic operation consequence: does the VMS legibility case use this corridor's design speed and scenario? |29| RLR-06 | Power/comms resilience: do battery runtime and network headroom clear the criteria using source-backed values? |30| RLR-07 | Comment and action closure: is every review comment closed, or carried with an owner and agreed action, or blocked by named missing data? |31| RLR-08 | Readiness consistency: does your final decision match your own matrix, findings, and action register? |32| RLR-09 | Claim boundary: does your review avoid unsupported approval, compliance, or acceptance claims? |33 34## Output35 36Write your complete review to `/workspace/output.md`. Explain your reasoning briefly in prose, then end with exactly one fenced JSON block:37 38```json39{40 "source_inventory": [{"doc_id": "...", "revision": "...", "status": "..."}],41 "identity_ledger": {42 "road_segment": "...",43 "low_point": "...",44 "low_point_chainage": "...",45 "datum": "...",46 "cabinet": "...",47 "vms": "...",48 "storm_case": "...",49 "network_case": "...",50 "battery": "..."51 },52 "review_matrix": {53 "RLR-01": {"status": "pass|fail|not_applicable|insufficient_data", "evidence": "..."},54 "RLR-02": {"status": "...", "evidence": "..."},55 "RLR-03": {"status": "...", "evidence": "..."},56 "RLR-04": {"status": "...", "evidence": "..."},57 "RLR-05": {"status": "...", "evidence": "..."},58 "RLR-06": {"status": "...", "evidence": "..."},59 "RLR-07": {"status": "...", "evidence": "..."},60 "RLR-08": {"status": "...", "evidence": "..."},61 "RLR-09": {"status": "...", "evidence": "..."}62 },63 "computed_evidence": {64 "peak_runoff_m3_s": 0.0,65 "gutter_approach_flow_m3_s": 0.0,66 "spread_width_m": 0.0,67 "allowable_spread_m": 0.0,68 "controlling_water_level_m": 0.0,69 "cabinet_freeboard_m": 0.0,70 "vms_message_margin_chars": 0.0,71 "battery_runtime_h": 0.0,72 "network_headroom_mbps": 0.073 },74 "findings": [75 {"item": "RLR-0X", "severity": "critical|minor", "source_id": "...", "object_id": "...", "consequence": "...", "action": "..."}76 ],77 "information_requests": [78 {"item": "RLR-0X", "missing_field": "...", "source_id": "..."}79 ],80 "action_register": [81 {"action": "...", "owner": "...", "linked_item": "RLR-0X"}82 ],83 "readiness_decision": "ready_to_issue|ready_with_carried_actions|not_ready_to_issue",84 "claim_boundary_statement": "..."85}86```87 88Rules for the structured block:89 90- `computed_evidence` values must come from your own recomputation from packet source values. Omit a key only when its inputs are missing from the packet (then raise the matching information request instead).91- Every `fail` needs at least one finding with a non-empty `source_id`, `object_id`, `consequence`, and `action`.92- Every `insufficient_data` needs an information request naming the exact missing field and its source document.93- Every `not_applicable` needs a scope reason in its matrix `evidence`.94- Carried actions must appear in `action_register` with an owner.95- `readiness_decision` must reconcile with your matrix: unresolved failures or missing critical evidence mean the package is not ready.96- `claim_boundary_statement` must state that this review covers a task-owned synthetic source packet and does not claim authority approval, accepted project evidence, full standards compliance, source-pack hardening, executable-verifier readiness, or benchmark readiness.97 Each generated task is drawn from one of these realistic scenario bands.
Site contexts ground each scenario in a real locale the model can use to infer hidden values.
urban_arterial
Urban arterial sag with kerbed gutter and dense ITS assets
suburban_collector
Suburban collector sag with roadside field cabinet
urban-arterial-sag-urban-arterial-preview — hard difficulty, some inputs hidden.
Urban arterial sag with kerbed gutter and dense ITS assets. urban-arterial-sag. Required outputs: rlr_01_status, rlr_02_status, rlr_03_status, rlr_04_status, rlr_05_status, rlr_06_status
Scenario context and visible inputs.
Executable tool: road-low-point-issue-review-package_calc.py
Inputs withheld at this difficulty.
Packet variant
packet_variant
Cabinet freeboard deficit
cabinet_freeboard_deficit_m
Cabinet freeboard margin
cabinet_freeboard_margin_m
Vms message margin chars target
vms_message_margin_chars_target
Network headroom margin mbps
network_headroom_margin_mbps
Hgl clearance target
hgl_clearance_target_m
Battery runtime margin h
battery_runtime_margin_h
Allowable spread margin
allowable_spread_margin_m
The scored JSON answer schema.
{
"rlr_01_status": <number>,
"rlr_02_status": <number>,
"rlr_03_status": <number>,
"rlr_04_status": <number>,
"rlr_05_status": <number>,
"rlr_06_status": <number>,
"rlr_07_status": <number>,
"rlr_08_status": <number>,
"rlr_09_status": <number>,
"readiness_code": <number>,
"required_findings_count": <number>,
"required_information_requests_count": <number>,
"required_carried_actions_count": <number>,
"peak_runoff_m3_s": <number>,
"gutter_approach_flow_m3_s": <number>,
"spread_width_m": <number>,
"allowable_spread_m": <number>,
"controlling_water_level_m": <number>,
"cabinet_freeboard_m": <number>,
"vms_message_margin_chars": <number>,
"battery_runtime_h": <number>,
"network_headroom_mbps": <number>
}