Rlr 01 status
rlr_01_status
Review matrix status code for RLR-01
Scores if within ±0% of the reference value.
Review-native long-horizon task over a generated SSC-15 product submittal issue packet. The agent receives a multi-file source packet in /workspace/sources/ (document register, submittal manifest, mill certificates, heat traceability table, product application schedule, deviation register, and criteria/comments), inventories it, preserves product, heat, certificate, and application identity, checks source-owned carbon-equivalent and property calculations as review evidence, assigns one status per review item, raises findings/information requests/actions, and issues a readiness decision. Packet variants inject missing heat traceability, stale certificate revisions, heat identity mismatches, copied standard applicability, open comments, or genuine carbon-equivalent failures. Scored by a stage-gated custom verifier, not per-key answer matching.
no-tool: The model must reason numerically unaided.
One template produces many comparable benchmark tasks while keeping the scoring contract fixed.
01
The reusable contract shown on this page.
02
An archetype and site context are sampled.
03
Inputs may be hidden at harder tiers.
04
The model responds with the declared outputs.
Inputs the model receives, and the outputs it is scored on.
8 inputs
Included directly in every task prompt.
Cev limit
cev_limit
Source-owned carbon-equivalent limit for the product use case
Required yield
required_yield_mpa
Source-owned minimum yield strength for the application grade
Required tensile
required_tensile_mpa
Source-owned minimum tensile strength for the application grade
Visible in easier tasks and withheld in one or more harder tiers.
Carbon equivalent margin
carbon_equivalent_margin
Derivation margin below the CEV limit for the controlling heat
Hidden at easy and medium and hard difficulty.
Carbon equivalent deficit
carbon_equivalent_deficit
Derivation exceedance above the CEV limit for the carbon-equivalent failure variant
Hidden at easy and medium and hard difficulty.
Yield strength margin
yield_strength_margin_mpa
Derivation margin between certificate yield and required yield
Hidden at easy and medium and hard difficulty.
Tensile strength margin
tensile_strength_margin_mpa
Derivation margin between certificate tensile strength and required tensile strength
Hidden at easy and medium and hard difficulty.
Packet variant
packet_variant
Hidden source-packet variant
Hidden at easy and medium and hard difficulty.
19 outputs
rlr_01_status
Review matrix status code for RLR-01
Scores if within ±0% of the reference value.
rlr_02_status
Review matrix status code for RLR-02
Scores if within ±0% of the reference value.
rlr_03_status
Review matrix status code for RLR-03
Scores if within ±0% of the reference value.
rlr_04_status
Review matrix status code for RLR-04
Scores if within ±0% of the reference value.
rlr_05_status
Review matrix status code for RLR-05
Scores if within ±0% of the reference value.
rlr_06_status
Review matrix status code for RLR-06
Scores if within ±0% of the reference value.
rlr_07_status
Review matrix status code for RLR-07
Scores if within ±0% of the reference value.
rlr_08_status
Review matrix status code for RLR-08
Scores if within ±0% of the reference value.
rlr_09_status
Review matrix status code for RLR-09
Scores if within ±0% of the reference value.
readiness_code
Readiness decision code
Scores if within ±0% of the reference value.
required_findings_count
Required finding count
Scores if within ±0% of the reference value.
required_information_requests_count
Required information request count
Scores if within ±0% of the reference value.
required_carried_actions_count
Required carried action count
Scores if within ±0% of the reference value.
carbon_equivalent_max
Maximum recomputed carbon equivalent for applied heats
Scores if within ±2% of the reference value.
carbon_equivalent_margin
CEV limit minus maximum recomputed carbon equivalent
Scores if within ±2% of the reference value.
yield_strength_margin_mpa
Minimum certified yield strength minus required yield strength
Scores if within ±2% of the reference value.
tensile_strength_margin_mpa
Minimum certified tensile strength minus required tensile strength
Scores if within ±2% of the reference value.
certificate_coverage_count
Applied heats covered by mill certificates
Scores if within ±0% of the reference value.
traceability_match_count
Application rows with traceable heat numbers
Scores if within ±0% of the reference value.
Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.
Some inputs hidden
Obvious packet defects: clean, missing heat traceability, or a carbon-equivalent failure
Hidden inputs
Packet variant restricted to: clean, missing_heat_number, carbon_equivalent_exceeds
Some inputs hidden
Full packet-variant distribution
Hidden inputs
Some inputs hidden
Subtle documentary defects: stale certificate revisions, heat identity mismatch, copied standard basis, carried comments
Hidden inputs
Packet variant restricted to: stale_certificate_revision, heat_number_mismatch, scenario_copy_forward, minor_open_comment_carried
The exact instruction and parameter contract used to generate this task, pinned to the published library source.
/workspace
Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.
1You are the independent reviewing engineer for a product submittal issue package covering steel product identity, mill certificates, heat traceability, application schedule, deviations, review comments, and criteria.2 3A source packet has been placed in `/workspace/sources/`. It contains a document register, submittal manifest, mill certificates, heat traceability table, product application schedule, deviation register, and a criteria memo with review comments. The packet is a task-owned synthetic source pack; treat it as the only source of numeric truth for this review.4 5Your job is not to redesign the product or accept the material for construction. Your job is to decide whether the package is ready to issue, and to produce an auditable review record.6 7## Review Workflow8 91. Inventory the source packet before drawing any conclusion. Record every document ID, revision, and status.102. Build an identity ledger: submittal, product, grade, heats, certificates, application schedule, deviation register, and criteria memo.113. Check for source conflicts, stale revisions, copied standards or use cases, missing evidence, heat/certificate mismatches, unresolved deviations, and open critical comments before accepting any package claim.124. Recompute the package's own checks only where they answer review items, using the assessment bases stated in the criteria memo. Do not import methods or values from outside the packet.135. Assign exactly one status to every review item: `pass`, `fail`, `not_applicable`, or `insufficient_data`.146. Do not invent missing values. Mark missing evidence as `insufficient_data` and request the exact missing field and source. A value that a source explicitly marks as pending or awaiting confirmation is missing evidence of this kind: it does not make otherwise-reconciling identifiers inconsistent and does not make evidence that is present untraceable; the check that cannot be completed without it takes `insufficient_data`.157. Convert every failure into a finding with a source pointer, affected object, consequence, and corrective action.168. Issue a readiness decision that reconciles with your matrix, findings, information requests, and action register.17 18## Review Matrix19 20Assess each item and give it exactly one status:21 22| Item | Review question |23|---|---|24| RLR-01 | Packet completeness: are all required submittal, certificate, traceability, application, deviation, and criteria files present with IDs and revisions? |25| RLR-02 | Object identity: do the submittal, product, grade, heat numbers, certificates, application rows, deviations, and criteria memo stay consistent? |26| RLR-03 | Certificate and property basis: are chemistry, mechanical properties, certificate scope, and source revisions traceable, current, and recomputable? |27| RLR-04 | Product compliance adequacy: do carbon equivalent, yield strength, tensile strength, certificate coverage, and traceability clear the source criteria for the same product use case? |28| RLR-05 | Application consequence: is the same product grade, standard, and application use case used across the schedule, certificates, deviations, and criteria? |29| RLR-06 | Source-status resilience: are certificate coverage and heat traceability source-backed for the issue package? |30| RLR-07 | Comment and action closure: is every critical comment closed, and are carried minor comments owner/action controlled? |31| RLR-08 | Readiness consistency: does your final decision match your own matrix, findings, information requests, and action register? |32| RLR-09 | Claim boundary: does your review avoid unsupported approval, compliance, source-hardening, executable-verifier, or benchmark-readiness claims? |33 34## Boundary Rules35 36- Use the review matrix definitions to decide the most specific affected item from the source packet. Do not infer a status from this instruction alone.37- If a source value needed for a recomputation is absent, omit the dependent `computed_evidence` key and raise an information request for the missing field and source. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.38- Assign additional failures only when the source packet gives independent evidence for them; do not double-count one source issue across unrelated matrix rows.39- RLR-08 passes when the readiness decision reconciles with the review matrix, findings, information requests, and action register.40- Every finding, information request, and action must name one exact RLR item. Use a single RLR item per register row; do not write combined items such as `RLR-04/RLR-06`.41- Do not rename computed_evidence keys.42 43## Output44 45Write your complete review to `/workspace/output.md`. Explain your reasoning briefly in prose, then end with exactly one fenced JSON block:46 47```json48{49 "source_inventory": [{"doc_id": "...", "revision": "...", "status": "..."}],50 "identity_ledger": {51 "submittal": "...",52 "product": "...",53 "grade": "...",54 "heat_a": "...",55 "heat_b": "...",56 "certificates": "...",57 "application_schedule": "...",58 "deviation_register": "...",59 "criteria_memo": "..."60 },61 "review_matrix": {62 "RLR-01": {"status": "pass|fail|not_applicable|insufficient_data", "evidence": "..."},63 "RLR-02": {"status": "...", "evidence": "..."},64 "RLR-03": {"status": "...", "evidence": "..."},65 "RLR-04": {"status": "...", "evidence": "..."},66 "RLR-05": {"status": "...", "evidence": "..."},67 "RLR-06": {"status": "...", "evidence": "..."},68 "RLR-07": {"status": "...", "evidence": "..."},69 "RLR-08": {"status": "...", "evidence": "..."},70 "RLR-09": {"status": "...", "evidence": "..."}71 },72 "computed_evidence": {73 "carbon_equivalent_max": 0.0,74 "carbon_equivalent_margin": 0.0,75 "yield_strength_margin_mpa": 0.0,76 "tensile_strength_margin_mpa": 0.0,77 "certificate_coverage_count": 0.0,78 "traceability_match_count": 0.079 },80 "findings": [81 {"item": "RLR-0X", "severity": "critical|minor", "source_id": "...", "object_id": "...", "consequence": "...", "action": "..."}82 ],83 "information_requests": [84 {"item": "RLR-0X", "missing_field": "...", "source_id": "..."}85 ],86 "action_register": [87 {"action": "...", "owner": "...", "linked_item": "RLR-0X"}88 ],89 "readiness_decision": "ready_to_issue|ready_with_carried_actions|not_ready_to_issue",90 "claim_boundary_statement": "..."91}92```93 94Rules for the structured block:95 96- `computed_evidence` values must come from your own recomputation from packet source values. Omit a key only when its inputs are missing from the packet, then raise the matching information request instead. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.97- Every `fail` needs at least one finding with a non-empty `source_id`, `object_id`, `consequence`, and `action`.98- Every `insufficient_data` needs an information request naming the exact missing field and its source document.99- Every `not_applicable` needs a scope reason in its matrix `evidence`.100- Carried actions must appear in `action_register` with an owner.101- `readiness_decision` must reconcile with your matrix: unresolved failures or missing critical evidence mean the package is not ready.102- `claim_boundary_statement` must state that this review covers a task-owned synthetic source packet and does not claim authority approval, accepted project evidence, full standards compliance, source-pack hardening, executable-verifier readiness, or benchmark readiness.103 Each generated task is drawn from one of these realistic scenario bands.
Site contexts ground each scenario in a real locale the model can use to infer hidden values.
bridge_plate_submittal
Bridge-plate steel product submittal with mill certificates and weldability criteria
plant_skid_frame_submittal
Plant skid-frame steel product submittal with heat traceability and deviation closeout
bridge-plate-product-review-bridge-plate-submittal-preview — hard difficulty, some inputs hidden.
Bridge-plate steel product submittal with mill certificates and weldability criteria. bridge-plate-product-review. Required outputs: rlr_01_status, rlr_02_status, rlr_03_status, rlr_04_status, rlr_05_status, rlr_06_status
Scenario context and visible inputs.
Executable tool: product-submittal-compliance-issue-review-package_calc.py
Inputs withheld at this difficulty.
Packet variant
packet_variant
Carbon equivalent deficit
carbon_equivalent_deficit
Yield strength margin mpa
yield_strength_margin_mpa
Carbon equivalent margin
carbon_equivalent_margin
Tensile strength margin mpa
tensile_strength_margin_mpa
The scored JSON answer schema.
{
"rlr_01_status": <number>,
"rlr_02_status": <number>,
"rlr_03_status": <number>,
"rlr_04_status": <number>,
"rlr_05_status": <number>,
"rlr_06_status": <number>,
"rlr_07_status": <number>,
"rlr_08_status": <number>,
"rlr_09_status": <number>,
"readiness_code": <number>,
"required_findings_count": <number>,
"required_information_requests_count": <number>,
"required_carried_actions_count": <number>,
"carbon_equivalent_max": <number>,
"carbon_equivalent_margin": <number>,
"yield_strength_margin_mpa": <number>,
"tensile_strength_margin_mpa": <number>,
"certificate_coverage_count": <number>,
"traceability_match_count": <number>
}