Rlr 01 status
rlr_01_status
Review matrix status code for RLR-01
Scores if within ±0% of the reference value.
Review-native long-horizon task over a generated SSC-06 pump station issue packet. The agent receives a multi-file source packet in /workspace/sources/ (document register, wet-well and suction geometry, rising-main schedule, pump curve and datasheet, motor and feeder schedule, duty operating case, criteria and comments), inventories it, preserves pump-station identity, checks the package's own calculations as review evidence, assigns one status per review item, raises findings/information requests/actions, and issues a readiness decision. Packet variants inject missing wet-well evidence, stale revisions, impeller identity mismatches, scenario copy-forward, open comments, or genuine NPSH criterion failures. Scored by a stage-gated custom verifier, not per-key answer matching.
no-tool: The model must reason numerically unaided.
Standards
One template produces many comparable benchmark tasks while keeping the scoring contract fixed.
01
The reusable contract shown on this page.
02
An archetype and site context are sampled.
03
Inputs may be hidden at harder tiers.
04
The model responds with the declared outputs.
Inputs the model receives, and the outputs it is scored on.
27 inputs
Included directly in every task prompt.
Design flow
design_flow_l_s
DUTY-06-CASE-01 design flow for PMP-06
Static lift
static_lift_m
DUTY-06-CASE-01 static lift from WW-06 to discharge boundary
Rising main length
rising_main_length_m
RM-06 rising-main length
Rising main diameter
rising_main_diameter_mm
RM-06 internal diameter
Hazen williams c
hazen_williams_c
RM-06 Hazen-Williams C factor
Minor loss coefficient
minor_loss_coefficient
RM-06 aggregate minor-loss coefficient
Fluid density
fluid_density_kg_m3
WW-06 source-owned fluid density
Pump efficiency pct
pump_efficiency_pct
PMP-06 pump efficiency at the selected duty point
Motor efficiency pct
motor_efficiency_pct
MOT-06 motor efficiency
Motor service factor
motor_service_factor
MOT-06 source-owned motor service factor
Atmospheric pressure
atmospheric_pressure_kpa_abs
WW-06 atmospheric pressure for NPSHa
Vapor pressure
vapor_pressure_kpa_abs
WW-06 vapor pressure for NPSHa
Wetwell min level above pump
wetwell_min_level_above_pump_m
WW-06 minimum operating level above pump centreline
Average level delta
average_level_delta_m
Difference between average and minimum wet-well operating levels
Suction loss
suction_loss_m
WW-06 suction-side loss at design duty
Minimum npsh margin
minimum_npsh_margin_m
CRIT-SSC06-001 minimum NPSH margin above NPSHr
Feeder voltage
feeder_voltage_v
FDR-06 nominal three-phase feeder voltage
Feeder length
feeder_length_km
FDR-06 feeder length
Feeder resistance ohm per km
feeder_resistance_ohm_per_km
FDR-06 conductor resistance
Feeder reactance ohm per km
feeder_reactance_ohm_per_km
FDR-06 conductor reactance
Motor power factor
motor_power_factor
MOT-06 source-owned motor power factor
Visible in easier tasks and withheld in one or more harder tiers.
Npsh margin target
npsh_margin_target_m
Derivation margin above the CRIT-SSC06-001 minimum NPSH margin
Hidden at easy and medium and hard difficulty.
Npsh margin deficit
npsh_margin_deficit_m
Derivation deficit below the CRIT-SSC06-001 minimum NPSH margin
Hidden at easy and medium and hard difficulty.
Pump head margin target
pump_head_margin_target_m
Derivation margin between PMP-06 curve head and total dynamic head
Hidden at easy and medium and hard difficulty.
Motor margin target
motor_margin_target_kw
Derivation margin between selected motor size and required motor power
Hidden at easy and medium and hard difficulty.
Voltage drop margin target percent
voltage_drop_margin_target_percent
Derivation margin between feeder voltage drop and maximum allowed drop
Hidden at easy and medium and hard difficulty.
Packet variant
packet_variant
Hidden source-packet variant
Hidden at easy and medium and hard difficulty.
21 outputs
rlr_01_status
Review matrix status code for RLR-01
Scores if within ±0% of the reference value.
rlr_02_status
Review matrix status code for RLR-02
Scores if within ±0% of the reference value.
rlr_03_status
Review matrix status code for RLR-03
Scores if within ±0% of the reference value.
rlr_04_status
Review matrix status code for RLR-04
Scores if within ±0% of the reference value.
rlr_05_status
Review matrix status code for RLR-05
Scores if within ±0% of the reference value.
rlr_06_status
Review matrix status code for RLR-06
Scores if within ±0% of the reference value.
rlr_07_status
Review matrix status code for RLR-07
Scores if within ±0% of the reference value.
rlr_08_status
Review matrix status code for RLR-08
Scores if within ±0% of the reference value.
rlr_09_status
Review matrix status code for RLR-09
Scores if within ±0% of the reference value.
readiness_code
Readiness decision code
Scores if within ±0% of the reference value.
required_findings_count
Required finding count
Scores if within ±0% of the reference value.
required_information_requests_count
Required information request count
Scores if within ±0% of the reference value.
required_carried_actions_count
Required carried action count
Scores if within ±0% of the reference value.
total_dynamic_head_m
Total dynamic head evidence
Scores if within ±2% of the reference value.
pump_head_margin_m
Pump curve head margin evidence
Scores if within ±2% of the reference value.
npsh_available_m
NPSH available evidence (absent when minimum wet-well level is missing)
Scores if within ±2% of the reference value.
npsh_margin_m
NPSH margin evidence (absent when minimum wet-well level is missing)
Scores if within ±2% of the reference value.
motor_input_kw
Motor input power evidence
Scores if within ±2% of the reference value.
motor_margin_kw
Selected motor margin evidence
Scores if within ±2% of the reference value.
feeder_voltage_drop_percent
Feeder voltage-drop evidence
Scores if within ±2% of the reference value.
voltage_drop_margin_percent
Voltage-drop margin evidence
Scores if within ±2% of the reference value.
Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.
Some inputs hidden
Obvious packet defects: clean, missing wet-well evidence, or a genuine NPSH failure
Hidden inputs
Packet variant restricted to: clean, missing_wetwell_min_level, npsh_margin_deficient
Some inputs hidden
Full packet-variant distribution
Hidden inputs
Some inputs hidden
Subtle documentary defects: stale revisions, identity mismatch, scenario copy-forward, carried comments
Hidden inputs
Packet variant restricted to: stale_pump_curve_revision, impeller_diameter_mismatch, scenario_copy_forward, minor_open_comment_carried
The exact instruction and parameter contract used to generate this task, pinned to the published library source.
/workspace
Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.
1You are the independent reviewing engineer for a pump station issue package covering wet-well geometry, suction basis, rising-main hydraulics, pump curve and datasheet, motor sizing, feeder voltage drop, duty case selection, and criteria/comments.2 3A source packet has been placed in `/workspace/sources/`. It contains a document register, wet-well and suction geometry, a rising-main schedule, a pump curve and datasheet extract, a motor and feeder schedule, a duty operating case, and a criteria memo with review comments. The packet is a task-owned synthetic source pack; treat it as the only source of numeric truth for this review.4 5Your job is not to redesign the pump station. Your job is to decide whether the package is ready to issue, and to produce an auditable review record.6 7## Review Workflow8 91. Inventory the source packet before drawing any conclusion. Record every document ID, revision, and status.102. Build an identity ledger: wet well, pump, impeller, rising main, motor, feeder, duty case, and criteria memo.113. Check for source conflicts, stale revisions, copied scenarios, missing evidence, pump/impeller mismatches, and open critical comments before accepting any package claim.124. Recompute the package's own calculations only where they answer review items, using the assessment bases stated in the criteria memo. Do not import methods or values from outside the packet.135. Assign exactly one status to every review item: `pass`, `fail`, `not_applicable`, or `insufficient_data`.146. Do not invent missing values. Mark missing evidence as `insufficient_data` and request the exact missing field and source. A value that a source explicitly marks as pending or awaiting confirmation is missing evidence of this kind: it does not make otherwise-reconciling identifiers inconsistent and does not make evidence that is present untraceable; the check that cannot be completed without it takes `insufficient_data`.157. Convert every failure into a finding with a source pointer, affected object, consequence, and corrective action.168. Issue a readiness decision that reconciles with your matrix, findings, information requests, and action register.17 18## Review Matrix19 20Assess each item and give it exactly one status:21 22| Item | Review question |23|---|---|24| RLR-01 | Packet completeness: are all required wet-well, rising-main, pump, motor/feeder, duty-case, and criteria files present with IDs and revisions? |25| RLR-02 | Object identity: do the pump ID, impeller, wet-well, rising main, motor, feeder, duty case, and criteria memo stay consistent across documents? |26| RLR-03 | Pump duty basis: are total dynamic head, pump head, motor input, and NPSH basis traceable, current, and recomputable? |27| RLR-04 | Pump adequacy: do pump head and NPSH margin clear the source criteria for the same pump and duty case? |28| RLR-05 | Scenario consequence: is the same duty flow and operating case used across wet well, pump curve, rising main, motor, and feeder documents? |29| RLR-06 | Secondary motor/feeder resilience: are motor sizing and feeder voltage drop source-backed and internally consistent with the selected duty point? |30| RLR-07 | Comment and action closure: is every critical comment closed, and are carried minor comments owner/action controlled? |31| RLR-08 | Readiness consistency: does your final decision match your own matrix, findings, information requests, and action register? |32| RLR-09 | Claim boundary: does your review avoid unsupported approval, compliance, source-hardening, executable-verifier, or benchmark-readiness claims? |33 34## Boundary Rules35 36- Use the review matrix definitions to decide the most specific affected item from the source packet. Do not infer a status from this instruction alone.37- If a source value needed for a recomputation is absent, omit the dependent `computed_evidence` key and raise an information request for the missing field and source. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.38- Assign additional failures only when the source packet gives independent evidence for them; do not double-count one source issue across unrelated matrix rows.39- RLR-08 passes when the readiness decision reconciles with the review matrix, findings, information requests, and action register.40- Every finding, information request, and action must name one exact RLR item. Use a single RLR item per register row; do not write combined items such as `RLR-04/RLR-06`.41- Do not rename computed_evidence keys.42 43## Output44 45Write your complete review to `/workspace/output.md`. Explain your reasoning briefly in prose, then end with exactly one fenced JSON block:46 47```json48{49 "source_inventory": [{"doc_id": "...", "revision": "...", "status": "..."}],50 "identity_ledger": {51 "wet_well": "...",52 "pump": "...",53 "impeller": "...",54 "rising_main": "...",55 "motor": "...",56 "feeder": "...",57 "duty_case": "...",58 "criteria_memo": "..."59 },60 "review_matrix": {61 "RLR-01": {"status": "pass|fail|not_applicable|insufficient_data", "evidence": "..."},62 "RLR-02": {"status": "...", "evidence": "..."},63 "RLR-03": {"status": "...", "evidence": "..."},64 "RLR-04": {"status": "...", "evidence": "..."},65 "RLR-05": {"status": "...", "evidence": "..."},66 "RLR-06": {"status": "...", "evidence": "..."},67 "RLR-07": {"status": "...", "evidence": "..."},68 "RLR-08": {"status": "...", "evidence": "..."},69 "RLR-09": {"status": "...", "evidence": "..."}70 },71 "computed_evidence": {72 "total_dynamic_head_m": 0.0,73 "pump_head_margin_m": 0.0,74 "npsh_available_m": 0.0,75 "npsh_margin_m": 0.0,76 "motor_input_kw": 0.0,77 "motor_margin_kw": 0.0,78 "feeder_voltage_drop_percent": 0.0,79 "voltage_drop_margin_percent": 0.080 },81 "findings": [82 {"item": "RLR-0X", "severity": "critical|minor", "source_id": "...", "object_id": "...", "consequence": "...", "action": "..."}83 ],84 "information_requests": [85 {"item": "RLR-0X", "missing_field": "...", "source_id": "..."}86 ],87 "action_register": [88 {"action": "...", "owner": "...", "linked_item": "RLR-0X"}89 ],90 "readiness_decision": "ready_to_issue|ready_with_carried_actions|not_ready_to_issue",91 "claim_boundary_statement": "..."92}93```94 95Rules for the structured block:96 97- `computed_evidence` values must come from your own recomputation from packet source values. Omit a key only when its inputs are missing from the packet, then raise the matching information request instead. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.98- Every `fail` needs at least one finding with a non-empty `source_id`, `object_id`, `consequence`, and `action`.99- Every `insufficient_data` needs an information request naming the exact missing field and its source document.100- Every `not_applicable` needs a scope reason in its matrix `evidence`.101- Carried actions must appear in `action_register` with an owner.102- `readiness_decision` must reconcile with your matrix: unresolved failures or missing critical evidence mean the package is not ready.103- `claim_boundary_statement` must state that this review covers a task-owned synthetic source packet and does not claim authority approval, accepted project evidence, full standards compliance, source-pack hardening, executable-verifier readiness, or benchmark readiness.104 Each generated task is drawn from one of these realistic scenario bands.
Site contexts ground each scenario in a real locale the model can use to infer hidden values.
municipal_wastewater_station
Municipal wet-well pump station with a moderate rising main
industrial_lift_station
Industrial lift station with longer feeder and higher duty head
municipal-wetwell-duty-municipal-wastewater-station-preview — hard difficulty, some inputs hidden.
Municipal wet-well pump station with a moderate rising main. municipal-wetwell-duty. Required outputs: rlr_01_status, rlr_02_status, rlr_03_status, rlr_04_status, rlr_05_status, rlr_06_status
Scenario context and visible inputs.
Executable tool: pump-station-duty-npsh-issue-review-package_calc.py
Inputs withheld at this difficulty.
Packet variant
packet_variant
Motor margin target
motor_margin_target_kw
Npsh margin deficit
npsh_margin_deficit_m
Voltage drop margin target
voltage_drop_margin_target_percent
Npsh margin target
npsh_margin_target_m
Pump head margin target
pump_head_margin_target_m
The scored JSON answer schema.
{
"rlr_01_status": <number>,
"rlr_02_status": <number>,
"rlr_03_status": <number>,
"rlr_04_status": <number>,
"rlr_05_status": <number>,
"rlr_06_status": <number>,
"rlr_07_status": <number>,
"rlr_08_status": <number>,
"rlr_09_status": <number>,
"readiness_code": <number>,
"required_findings_count": <number>,
"required_information_requests_count": <number>,
"required_carried_actions_count": <number>,
"total_dynamic_head_m": <number>,
"pump_head_margin_m": <number>,
"npsh_available_m": <number>,
"npsh_margin_m": <number>,
"motor_input_kw": <number>,
"motor_margin_kw": <number>,
"feeder_voltage_drop_percent": <number>,
"voltage_drop_margin_percent": <number>
}