Rlr 01 status
rlr_01_status
Packet completeness status code (0 pass, 1 fail, 2 not applicable, 3 insufficient data)
Scores if within ±0% of the reference value.
Review-native long-horizon task over a generated SSC-01 multimodal corridor comment-response packet. The agent receives a multi-file source packet in /workspace/sources/ (document register, comment register, marked-up plan, drainage recalculation, signal and pedestrian recalculation, VMS operations note, electrical feeder check, and criteria/comments), inventories it, preserves corridor identity, recomputes the package's own chainage, drainage, pedestrian, VMS, voltage-drop, and closeout evidence, assigns one status per review item, raises findings/information requests/actions, and issues a readiness decision. Packet variants inject missing revised chainage evidence, stale change-register revisions, chainage identity mismatch, scenario copy-forward, open comments, carried comments, or a genuine unsupported downstream repair. Scored by a stage-gated custom verifier, not per-key answer matching.
no-tool: The model must reason numerically unaided.
One template produces many comparable benchmark tasks while keeping the scoring contract fixed.
01
The reusable contract shown on this page.
02
An archetype and site context are sampled.
03
Inputs may be hidden at harder tiers.
04
The model responds with the declared outputs.
Inputs the model receives, and the outputs it is scored on.
23 inputs
Included directly in every task prompt.
Original chainage
original_chainage_m
C-008 original comment chainage on RD-SSC01-001
Chainage delta
chainage_delta_m
Derived comment chainage movement from markup to revised package
Revised hgl m
revised_hgl_m
DRN-08 revised HGL for the comment-response location
Minimum hgl clearance
minimum_hgl_clearance_mm
Criteria minimum road level to HGL clearance
Pedestrian startup time
pedestrian_startup_time_s
SG-08 pedestrian startup allowance
Revised crossing width
revised_crossing_width_m
SG-08 revised pedestrian crossing width
Pedestrian walk speed
pedestrian_walk_speed_m_s
SG-08 pedestrian walking speed
Vms character height
vms_character_height_in
VMS-08 revised character height
Approach speed kmh
approach_speed_kmh
CASE-08 approach speed used for VMS reading-time check
Reading rate
reading_rate_chars_s
CASE-08 accepted VMS reading rate
Revised device load
revised_device_load_w
FEED-08 revised connected load after comment-response changes
Feeder length
feeder_length_km
FEED-08 route length
Conductor resistance
conductor_resistance_ohm_km
FEED-08 conductor resistance
Feeder voltage
feeder_voltage_v
FEED-08 nominal feeder voltage
Power factor
power_factor
FEED-08 load power factor
Review comments total
review_comments_total
CMT-SSC01-008 total comments in the response batch
Impacted calculation count
impacted_calculation_count
CMT-SSC01-008 count of calculations touched by the review change
Visible in easier tasks and withheld in one or more harder tiers.
Hgl clearance margin
hgl_clearance_margin_mm
Hidden derivation margin between revised clearance and required HGL clearance
Hidden at easy and medium and hard difficulty.
Ped clearance margin
ped_clearance_margin_s
Hidden derivation margin between available and required pedestrian clearance
Hidden at easy and medium and hard difficulty.
Vms message margin
vms_message_margin_chars
Hidden derivation margin between readable VMS capacity and message length
Hidden at easy and medium and hard difficulty.
Voltage drop margin percent
voltage_drop_margin_percent
Hidden derivation margin between computed and allowable voltage drop
Hidden at easy and medium and hard difficulty.
Voltage drop deficit percent
voltage_drop_deficit_percent
Hidden derivation deficit for the unsupported-downstream-repair variant
Hidden at easy and medium and hard difficulty.
Packet variant
packet_variant
Hidden packet-defect variant controlling gold review statuses
Hidden at easy and medium and hard difficulty.
24 outputs
rlr_01_status
Packet completeness status code (0 pass, 1 fail, 2 not applicable, 3 insufficient data)
Scores if within ±0% of the reference value.
rlr_02_status
Object identity status code
Scores if within ±0% of the reference value.
rlr_03_status
Review-response basis status code
Scores if within ±0% of the reference value.
rlr_04_status
Review-response adequacy status code
Scores if within ±0% of the reference value.
rlr_05_status
Scenario consequence status code
Scores if within ±0% of the reference value.
rlr_06_status
Secondary-discipline resilience status code
Scores if within ±0% of the reference value.
rlr_07_status
Comment and action closure status code
Scores if within ±0% of the reference value.
rlr_08_status
Readiness decision consistency status code
Scores if within ±0% of the reference value.
rlr_09_status
Claim boundary status code
Scores if within ±0% of the reference value.
readiness_code
Gold readiness decision (0 ready, 1 ready with carried actions, 2 not ready)
Scores if within ±0% of the reference value.
required_findings_count
Findings the review must raise
Scores if within ±0% of the reference value.
required_information_requests_count
Information requests the review must raise
Scores if within ±0% of the reference value.
required_carried_actions_count
Carried actions the review must record
Scores if within ±0% of the reference value.
changed_chainage_delta_m
Revised comment chainage minus original comment chainage
Scores if within ±2% of the reference value.
hgl_clearance_mm
Revised road level minus revised HGL
Scores if within ±2% of the reference value.
hgl_clearance_margin_mm
HGL clearance minus minimum required clearance
Scores if within ±2% of the reference value.
ped_clearance_required_s
Required pedestrian clearance time
Scores if within ±2% of the reference value.
ped_clearance_margin_s
Available pedestrian clearance minus required clearance
Scores if within ±2% of the reference value.
vms_reading_time_s
Available VMS reading time
Scores if within ±2% of the reference value.
vms_message_margin_chars
Readable character capacity minus revised message length
Scores if within ±2% of the reference value.
feeder_voltage_drop_percent
Voltage drop at revised device load
Scores if within ±2% of the reference value.
voltage_drop_margin_percent
Allowable voltage drop minus computed voltage drop
Scores if within ±2% of the reference value.
comment_closeout_percent
Closed comments divided by total comments
Scores if within ±2% of the reference value.
impacted_calculation_count
Number of discipline calculations impacted by the review change
Scores if within ±2% of the reference value.
Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.
Some inputs hidden
Obvious packet defects: clean, missing revised chainage, or a genuine downstream-repair failure
Hidden inputs
Packet variant restricted to: clean, missing_revised_chainage, unsupported_downstream_repair
Some inputs hidden
Full packet-variant distribution
Hidden inputs
Some inputs hidden
Subtle documentary defects: stale register, identity drift, copied scenario, and carried comments
Hidden inputs
Packet variant restricted to: stale_change_register_revision, chainage_identity_mismatch, scenario_copy_forward, minor_open_comment_carried
The exact instruction and parameter contract used to generate this task, pinned to the published library source.
/workspace
Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.
1You are the independent reviewing engineer for a multimodal corridor comment-response issue package covering one road corridor, one review comment, one drainage recalculation, one signal and pedestrian recalculation, one VMS operation note, and one ITS feeder check.2 3A source packet has been placed in `/workspace/sources/`. It contains a document register, comment register, marked-up plan and long section, drainage recalculation, signal and pedestrian recalculation, VMS operations note, electrical feeder check, and a criteria memo. The packet is a task-owned synthetic source pack; treat it as the only source of numeric truth for this review.4 5Your job is not to redesign the corridor. Your job is to decide whether the comment-response package is ready to issue, and to produce an auditable review record.6 7## Review Workflow8 91. Inventory the source packet before drawing any conclusion. Record every document ID, revision, and status.102. Build an identity ledger: corridor, comment, original and revised chainage, datum, scenario, drainage object, signal group, VMS device, and field feeder.113. Check for source conflicts, stale revisions, chainage drift, copied scenarios, missing evidence, and open critical comments before accepting any package claim.124. Recompute the package's own calculations only where they answer review items, using the assessment bases stated in the criteria memo. Do not import methods or values from outside the packet.135. Assign exactly one status to every review item: `pass`, `fail`, `not_applicable`, or `insufficient_data`.146. Do not invent missing values. Mark missing evidence as `insufficient_data` and request the exact missing field and source. A value that a source explicitly marks as pending or awaiting confirmation is missing evidence of this kind: it does not make otherwise-reconciling identifiers inconsistent and does not make evidence that is present untraceable; the check that cannot be completed without it takes `insufficient_data`.157. Convert every failure into a finding with a source pointer, affected object, consequence, and corrective action.168. Issue a readiness decision that reconciles with your matrix, findings, and action register.17 18## Review Discipline19 20Use the most specific review item for each defect based on the source packet and the review matrix definitions. Do not infer a status from this instruction alone.21 22- If a source value needed for a recomputation is absent, omit the dependent `computed_evidence` key and raise an information request for the missing field and source. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.23- Assign additional failures only when the source packet gives independent evidence for them; do not double-count one source issue across unrelated matrix rows.24- RLR-08 is reviewer self-consistency: it passes when your readiness decision reconciles with your matrix, findings, information requests, and action register.25- Every finding, information request, and action must name one exact RLR item.26 27## Review Matrix28 29Assess each item and give it exactly one status:30 31| Item | Review question |32|---|---|33| RLR-01 | Packet completeness: are all required source documents present with IDs and revisions? |34| RLR-02 | Object identity: do corridor, comment, chainage, datum, scenario, drainage, signal, VMS, and feeder identities stay consistent across documents? |35| RLR-03 | Review-response basis: are the change ledger and impacted calculation count traceable to current source revisions and recomputable evidence? |36| RLR-04 | Review-response adequacy: does the changed chainage/scenario propagate through HGL, pedestrian, VMS, voltage-drop, and comment-closeout checks? |37| RLR-05 | Scenario consequence: is the same corridor scenario used across disciplines rather than copied from another corridor? |38| RLR-06 | Secondary-discipline resilience: are VMS legibility and feeder voltage-drop checks source-backed and internally consistent? |39| RLR-07 | Comment and action closure: is every review comment closed, or carried with an owner and agreed action, or blocked by named missing data? |40| RLR-08 | Readiness consistency: does your final decision match your own matrix, findings, and action register? |41| RLR-09 | Claim boundary: does your review avoid unsupported approval, compliance, or acceptance claims? |42 43## Output44 45Write your complete review to `/workspace/output.md`. Explain your reasoning briefly in prose, then end with exactly one fenced JSON block:46 47```json48{49 "source_inventory": [{"doc_id": "...", "revision": "...", "status": "..."}],50 "identity_ledger": {51 "corridor": "...",52 "comment": "...",53 "original_chainage": "...",54 "revised_chainage": "...",55 "datum": "...",56 "scenario": "...",57 "drainage_object": "...",58 "signal_group": "...",59 "vms_device": "...",60 "field_feeder": "..."61 },62 "review_matrix": {63 "RLR-01": {"status": "pass|fail|not_applicable|insufficient_data", "evidence": "..."},64 "RLR-02": {"status": "...", "evidence": "..."},65 "RLR-03": {"status": "...", "evidence": "..."},66 "RLR-04": {"status": "...", "evidence": "..."},67 "RLR-05": {"status": "...", "evidence": "..."},68 "RLR-06": {"status": "...", "evidence": "..."},69 "RLR-07": {"status": "...", "evidence": "..."},70 "RLR-08": {"status": "...", "evidence": "..."},71 "RLR-09": {"status": "...", "evidence": "..."}72 },73 "computed_evidence": {74 "changed_chainage_delta_m": 0.0,75 "hgl_clearance_mm": 0.0,76 "hgl_clearance_margin_mm": 0.0,77 "ped_clearance_required_s": 0.0,78 "ped_clearance_margin_s": 0.0,79 "vms_reading_time_s": 0.0,80 "vms_message_margin_chars": 0.0,81 "feeder_voltage_drop_percent": 0.0,82 "voltage_drop_margin_percent": 0.0,83 "comment_closeout_percent": 0.0,84 "impacted_calculation_count": 0.085 },86 "findings": [87 {"item": "RLR-0X", "severity": "critical|minor", "source_id": "...", "object_id": "...", "consequence": "...", "action": "..."}88 ],89 "information_requests": [90 {"item": "RLR-0X", "missing_field": "...", "source_id": "..."}91 ],92 "action_register": [93 {"action": "...", "owner": "...", "linked_item": "RLR-0X"}94 ],95 "readiness_decision": "ready_to_issue|ready_with_carried_actions|not_ready_to_issue",96 "claim_boundary_statement": "..."97}98```99 100Rules for the structured block:101 102- Do not rename computed_evidence keys. Use the exact key names shown in the schema. Use `voltage_drop_margin_percent`, not `feeder_voltage_drop_margin_percent`.103- `computed_evidence` values must come from your own recomputation from packet source values. Omit a key only when its inputs are missing from the packet, then raise the matching information request instead. Do not include missing or unrecomputable evidence keys with `null`, `0`, or placeholder values.104- Every `fail` needs at least one finding with a non-empty `source_id`, `object_id`, `consequence`, and `action`. Each `findings[].item` value must be a single RLR item such as `RLR-05`; do not combine multiple IDs in one item field.105- Every `insufficient_data` needs an information request naming the exact missing field and its source document. Each `information_requests[].item` value must be a single RLR item such as `RLR-04`.106- Every `not_applicable` needs a scope reason in its matrix `evidence`.107- Carried actions must appear in `action_register` with an owner. Each `action_register[].linked_item` value must be a single RLR item.108- `readiness_decision` must reconcile with your matrix: unresolved failures or missing critical evidence mean the package is not ready.109- `claim_boundary_statement` must state that this review covers a task-owned synthetic source packet and does not claim authority approval, accepted project evidence, full standards compliance, source-pack hardening, executable-verifier readiness, or benchmark readiness.110 Each generated task is drawn from one of these realistic scenario bands.
Site contexts ground each scenario in a real locale the model can use to infer hidden values.
urban_arterial_response
Urban arterial comment-response package with wider crossing and larger ITS load
suburban_collector_response
Suburban collector comment-response package with moderate crossing width and shorter feeder
urban-arterial-comment-response-urban-arterial-response-preview — hard difficulty, some inputs hidden.
Urban arterial comment-response package with wider crossing and larger ITS load. urban-arterial-comment-response. Required outputs: rlr_01_status, rlr_02_status, rlr_03_status, rlr_04_status, rlr_05_status, rlr_06_status
Scenario context and visible inputs.
Executable tool: corridor-comment-response-issue-review-package_calc.py
Inputs withheld at this difficulty.
Ped clearance margin s
ped_clearance_margin_s
Packet variant
packet_variant
Voltage drop deficit
voltage_drop_deficit_percent
Vms message margin chars
vms_message_margin_chars
Voltage drop margin
voltage_drop_margin_percent
Hgl clearance margin
hgl_clearance_margin_mm
The scored JSON answer schema.
{
"rlr_01_status": <number>,
"rlr_02_status": <number>,
"rlr_03_status": <number>,
"rlr_04_status": <number>,
"rlr_05_status": <number>,
"rlr_06_status": <number>,
"rlr_07_status": <number>,
"rlr_08_status": <number>,
"rlr_09_status": <number>,
"readiness_code": <number>,
"required_findings_count": <number>,
"required_information_requests_count": <number>,
"required_carried_actions_count": <number>,
"changed_chainage_delta_m": <number>,
"hgl_clearance_mm": <number>,
"hgl_clearance_margin_mm": <number>,
"ped_clearance_required_s": <number>,
"ped_clearance_margin_s": <number>,
"vms_reading_time_s": <number>,
"vms_message_margin_chars": <number>,
"feeder_voltage_drop_percent": <number>,
"voltage_drop_margin_percent": <number>,
"comment_closeout_percent": <number>,
"impacted_calculation_count": <number>
}