LIVEdataset aec-bench@releasetasks 552models 18last submission · built
mechanicalno-tool

Fire Water Storage Hazard Issue Review Package

Review-native long-horizon task over a generated SSC-19 fire-water issue packet. The agent receives a multi-file source packet in /workspace/sources/ (document register, hazard and storage arrangement, sprinkler/hydrant demand sheet, tank and pump schedule, fire strategy case, water-supply basis, criteria and comments), inventories it, preserves hazard and asset identity, checks source-owned classification and storage calculations as review evidence, assigns one status per review item, raises findings/information requests/actions, and issues a readiness decision. Packet variants inject missing hazard evidence, stale revisions, tank identity mismatches, scenario copy-forward, open comments, or genuine storage failures under the true hazard class. Scored by a stage-gated custom verifier, not per-key answer matching.

no-tool: The model must reason numerically unaided.

How this task is generated

One template produces many comparable benchmark tasks while keeping the scoring contract fixed.

  1. 01

    Template

    The reusable contract shown on this page.

  2. 02

    Scenario

    An archetype and site context are sampled.

  3. 03

    Difficulty tier

    Inputs may be hidden at harder tiers.

  4. 04

    Task prompt

    The model responds with the declared outputs.

Parameters

Inputs the model receives, and the outputs it is scored on.

Inputs

7 inputs

Always given

Included directly in every task prompt.

2
  • Storage height

    storage_height_m

    HAZ-19 storage height for the commodity arrangement

    6 – 8.5 m
  • Remote area adjustment

    remote_area_adjustment_m2

    Source-owned remote-area adjustment applied to the class baseline

    -12 – 18 m2

Hidden at higher difficulty

Visible in easier tasks and withheld in one or more harder tiers.

5
  • Storage margin

    storage_margin_m3

    Derivation margin between true required volume and available tank storage

    Hidden at easy and medium and hard difficulty.

    20 – 80 m3
  • Storage deficit

    storage_deficit_m3

    Derivation deficit below true required storage for the storage-deficient variant

    Hidden at easy and medium and hard difficulty.

    8 – 35 m3
  • Pump capacity margin

    pump_capacity_margin_l_min

    Derivation margin between pump rated capacity and total fire-water demand

    Hidden at easy and medium and hard difficulty.

    250 – 900 L/min
  • Tank volume mismatch delta

    tank_volume_mismatch_delta_m3

    Documentary mismatch between tank schedule and demand-sheet tank volume

    Hidden at easy and medium and hard difficulty.

    35 – 90 m3
  • Packet variant

    packet_variant

    Hidden source-packet variant

    Hidden at easy and medium and hard difficulty.

    cleanmissing_commodity_classificationstale_hazard_basis_revisiontank_volume_mismatchscenario_copy_forwardopen_critical_commentminor_open_comment_carriedstorage_deficient_under_true_class

Scored outputs

21 outputs

Rlr 01 status

rlr_01_status

Review matrix status code for RLR-01

Scores if within ±0% of the reference value.

Rlr 02 status

rlr_02_status

Review matrix status code for RLR-02

Scores if within ±0% of the reference value.

Rlr 03 status

rlr_03_status

Review matrix status code for RLR-03

Scores if within ±0% of the reference value.

Rlr 04 status

rlr_04_status

Review matrix status code for RLR-04

Scores if within ±0% of the reference value.

Rlr 05 status

rlr_05_status

Review matrix status code for RLR-05

Scores if within ±0% of the reference value.

Rlr 06 status

rlr_06_status

Review matrix status code for RLR-06

Scores if within ±0% of the reference value.

Rlr 07 status

rlr_07_status

Review matrix status code for RLR-07

Scores if within ±0% of the reference value.

Rlr 08 status

rlr_08_status

Review matrix status code for RLR-08

Scores if within ±0% of the reference value.

Rlr 09 status

rlr_09_status

Review matrix status code for RLR-09

Scores if within ±0% of the reference value.

Readiness code

readiness_code

Readiness decision code

Scores if within ±0% of the reference value.

Required findings count

required_findings_count

Required finding count

Scores if within ±0% of the reference value.

Required information requests count

required_information_requests_count

Required information request count

Scores if within ±0% of the reference value.

Required carried actions count

required_carried_actions_count

Required carried action count

Scores if within ±0% of the reference value.

Design density mm min

design_density_mm_min

True source-owned sprinkler design density

Scores if within ±2% of the reference value.

Design area m2

design_area_m2

True source-owned sprinkler design area

Scores if within ±2% of the reference value.

Sprinkler demand l min

sprinkler_demand_l_min

Sprinkler density-area demand

Scores if within ±2% of the reference value.

Hose allowance l min

hose_allowance_l_min

Hose allowance

Scores if within ±2% of the reference value.

Required duration min

required_duration_min

Required fire-water duration

Scores if within ±2% of the reference value.

Required volume m3

required_volume_m3

Required fire-water storage volume

Scores if within ±2% of the reference value.

Storage volume margin m3

storage_volume_margin_m3

Available storage minus required storage

Scores if within ±2% of the reference value.

Pump capacity margin l min

pump_capacity_margin_l_min

Fire pump rated capacity minus total fire-water demand

Scores if within ±2% of the reference value.

Difficulty

Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.

easy

Some inputs hidden

Obvious packet defects: clean, missing classification evidence, or a genuine storage failure

Hidden inputs

  • Packet variantpacket_variant
  • Storage margin m3storage_margin_m3
  • Storage deficit m3storage_deficit_m3
  • Pump capacity margin l minpump_capacity_margin_l_min
  • Tank volume mismatch delta m3tank_volume_mismatch_delta_m3

Packet variant restricted to: clean, missing_commodity_classification, storage_deficient_under_true_class

medium

Some inputs hidden

Full packet-variant distribution

Hidden inputs

  • Packet variantpacket_variant
  • Storage margin m3storage_margin_m3
  • Storage deficit m3storage_deficit_m3
  • Pump capacity margin l minpump_capacity_margin_l_min
  • Tank volume mismatch delta m3tank_volume_mismatch_delta_m3

hard

Some inputs hidden

Subtle documentary defects: stale revisions, tank identity mismatch, scenario copy-forward, carried comments

Hidden inputs

  • Packet variantpacket_variant
  • Storage margin m3storage_margin_m3
  • Storage deficit m3storage_deficit_m3
  • Pump capacity margin l minpump_capacity_margin_l_min
  • Tank volume mismatch delta m3tank_volume_mismatch_delta_m3

Packet variant restricted to: stale_hazard_basis_revision, tank_volume_mismatch, scenario_copy_forward, minor_open_comment_carried

Task bundle

The exact instruction and parameter contract used to generate this task, pinned to the published library source.

/workspace

  • instruction.md

Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.

1You are the independent reviewing engineer for a fire-water issue package covering hazard classification, storage arrangement, sprinkler demand, hydrant and pump capacity, fire-water storage, AHJ case selection, and criteria/comments.2 3A source packet has been placed in `/workspace/sources/`. It contains a document register, hazard and storage arrangement, sprinkler/hydrant demand sheet, tank and pump schedule, fire strategy operating case, water-supply basis, and a criteria memo with review comments. The packet is a task-owned synthetic source pack; treat it as the only source of numeric truth for this review.4 5Your job is not to redesign the fire-protection system. Your job is to decide whether the package is ready to issue, and to produce an auditable review record.6 7## Review Workflow8 91. Inventory the source packet before drawing any conclusion. Record every document ID, revision, and status.102. Build an identity ledger: building, hazard arrangement, commodity evidence, sprinkler system, tank, pump, water-supply basis, fire strategy case, and criteria memo.113. Check for source conflicts, stale revisions, copied scenarios, missing evidence, tank/pump mismatches, and open critical comments before accepting any package claim.124. Recompute the package's own calculations only where they answer review items, using the assessment bases stated in the criteria memo. Do not import methods or values from outside the packet.135. Assign exactly one status to every review item: `pass`, `fail`, `not_applicable`, or `insufficient_data`.146. Do not invent missing values. Mark missing evidence as `insufficient_data` and request the exact missing field and source. A value that a source explicitly marks as pending or awaiting confirmation is missing evidence of this kind: it does not make otherwise-reconciling identifiers inconsistent and does not make evidence that is present untraceable; the check that cannot be completed without it takes `insufficient_data`.157. Convert every failure into a finding with a source pointer, affected object, consequence, and corrective action.168. Issue a readiness decision that reconciles with your matrix, findings, information requests, and action register.17 18## Review Matrix19 20Assess each item and give it exactly one status:21 22| Item | Review question |23|---|---|24| RLR-01 | Packet completeness: are all required hazard, demand, tank/pump, fire strategy, water-supply, and criteria files present with IDs and revisions? |25| RLR-02 | Object identity: do the building, hazard arrangement, sprinkler system, tank, pump, water-supply basis, fire scenario, and criteria memo stay consistent? |26| RLR-03 | Hazard and demand basis: are hazard classification, design density, design area, hose allowance, duration, and demand traceable, current, and recomputable? |27| RLR-04 | Storage and pump adequacy: do required storage volume and fire pump capacity clear the source criteria for the same hazard and scenario? |28| RLR-05 | Scenario consequence: is the same fire strategy case and storage arrangement used across hazard, demand, tank, pump, and water-supply documents? |29| RLR-06 | Secondary water-supply resilience: is pump capacity source-backed and internally consistent with the selected demand case? |30| RLR-07 | Comment and action closure: is every critical comment closed, and are carried minor comments owner/action controlled? |31| RLR-08 | Readiness consistency: does your final decision match your own matrix, findings, information requests, and action register? |32| RLR-09 | Claim boundary: does your review avoid unsupported approval, compliance, source-hardening, executable-verifier, or benchmark-readiness claims? |33 34## Boundary Rules35 36- Use the review matrix definitions to decide the most specific affected item from the source packet. Do not infer a status from this instruction alone.37- If a source value needed for a recomputation is absent, omit the dependent `computed_evidence` key and raise an information request for the missing field and source. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.38- Assign additional failures only when the source packet gives independent evidence for them; do not double-count one source issue across unrelated matrix rows.39- RLR-08 passes when the readiness decision reconciles with the review matrix, findings, information requests, and action register.40- Every finding, information request, and action must name one exact RLR item. Use a single RLR item per register row; do not write combined items such as `RLR-04/RLR-06`.41- Do not rename computed_evidence keys.42 43## Output44 45Write your complete review to `/workspace/output.md`. Explain your reasoning briefly in prose, then end with exactly one fenced JSON block:46 47```json48{49 "source_inventory": [{"doc_id": "...", "revision": "...", "status": "..."}],50 "identity_ledger": {51 "building": "...",52 "hazard_arrangement": "...",53 "commodity_evidence": "...",54 "sprinkler_system": "...",55 "tank": "...",56 "pump": "...",57 "water_supply_basis": "...",58 "fire_strategy_case": "...",59 "criteria_memo": "..."60 },61 "review_matrix": {62 "RLR-01": {"status": "pass|fail|not_applicable|insufficient_data", "evidence": "..."},63 "RLR-02": {"status": "...", "evidence": "..."},64 "RLR-03": {"status": "...", "evidence": "..."},65 "RLR-04": {"status": "...", "evidence": "..."},66 "RLR-05": {"status": "...", "evidence": "..."},67 "RLR-06": {"status": "...", "evidence": "..."},68 "RLR-07": {"status": "...", "evidence": "..."},69 "RLR-08": {"status": "...", "evidence": "..."},70 "RLR-09": {"status": "...", "evidence": "..."}71 },72 "computed_evidence": {73 "design_density_mm_min": 0.0,74 "design_area_m2": 0.0,75 "sprinkler_demand_l_min": 0.0,76 "hose_allowance_l_min": 0.0,77 "required_duration_min": 0.0,78 "required_volume_m3": 0.0,79 "storage_volume_margin_m3": 0.0,80 "pump_capacity_margin_l_min": 0.081 },82 "findings": [83 {"item": "RLR-0X", "severity": "critical|minor", "source_id": "...", "object_id": "...", "consequence": "...", "action": "..."}84 ],85 "information_requests": [86 {"item": "RLR-0X", "missing_field": "...", "source_id": "..."}87 ],88 "action_register": [89 {"action": "...", "owner": "...", "linked_item": "RLR-0X"}90 ],91 "readiness_decision": "ready_to_issue|ready_with_carried_actions|not_ready_to_issue",92 "claim_boundary_statement": "..."93}94```95 96Rules for the structured block:97 98- `computed_evidence` values must come from your own recomputation from packet source values. Omit a key only when its inputs are missing from the packet, then raise the matching information request instead. Do not include missing or unrecomputable keys with `null`, `0`, or placeholder values.99- Every `fail` needs at least one finding with a non-empty `source_id`, `object_id`, `consequence`, and `action`.100- Every `insufficient_data` needs an information request naming the exact missing field and its source document.101- Every `not_applicable` needs a scope reason in its matrix `evidence`.102- Carried actions must appear in `action_register` with an owner.103- `readiness_decision` must reconcile with your matrix: unresolved failures or missing critical evidence mean the package is not ready.104- `claim_boundary_statement` must state that this review covers a task-owned synthetic source packet and does not claim authority approval, accepted project evidence, full standards compliance, source-pack hardening, executable-verifier readiness, or benchmark readiness.105

Scenario archetypes

Each generated task is drawn from one of these realistic scenario bands.

Site contexts ground each scenario in a real locale the model can use to infer hidden values.

High bay warehouse

high_bay_warehouse

High-bay warehouse storage arrangement with rack storage and AHJ/FM criteria

high-bay-warehouse-fire-waterrack-storage-hazard-review
Parameter ranges
storage_height_m
6.8 – 8.5
remote_area_adjustment_m2
-8 – 18

Medium height storage

medium_height_storage

Medium-height storage building with racked commodity and fire-water tank

medium-height-storage-fire-waterdistribution-building-hazard-review
Parameter ranges
storage_height_m
6 – 6.7
remote_area_adjustment_m2
-12 – 10

Example task

high-bay-warehouse-fire-water-high-bay-warehouse-previewhard difficulty, some inputs hidden.

High-bay warehouse storage arrangement with rack storage and AHJ/FM criteria. high-bay-warehouse-fire-water. Required outputs: rlr_01_status, rlr_02_status, rlr_03_status, rlr_04_status, rlr_05_status, rlr_06_status

The model sees

Scenario context and visible inputs.

storage_height_m
6.8 to 8.5 m
remote_area_adjustment_m2
-8 to 18 m2

Executable tool: fire-water-storage-hazard-issue-review-package_calc.py

The model must infer

Inputs withheld at this difficulty.

  • Pump capacity margin l min

    pump_capacity_margin_l_min

  • Storage margin m3

    storage_margin_m3

  • Packet variant

    packet_variant

  • Storage deficit m3

    storage_deficit_m3

  • Tank volume mismatch delta m3

    tank_volume_mismatch_delta_m3

The model must produce

The scored JSON answer schema.

{
  "rlr_01_status": <number>,
  "rlr_02_status": <number>,
  "rlr_03_status": <number>,
  "rlr_04_status": <number>,
  "rlr_05_status": <number>,
  "rlr_06_status": <number>,
  "rlr_07_status": <number>,
  "rlr_08_status": <number>,
  "rlr_09_status": <number>,
  "readiness_code": <number>,
  "required_findings_count": <number>,
  "required_information_requests_count": <number>,
  "required_carried_actions_count": <number>,
  "design_density_mm_min": <number>,
  "design_area_m2": <number>,
  "sprinkler_demand_l_min": <number>,
  "hose_allowance_l_min": <number>,
  "required_duration_min": <number>,
  "required_volume_m3": <number>,
  "storage_volume_margin_m3": <number>,
  "pump_capacity_margin_l_min": <number>
}
  • rlr_01_status · scored within ±0%
  • rlr_02_status · scored within ±0%
  • rlr_03_status · scored within ±0%
  • rlr_04_status · scored within ±0%
  • rlr_05_status · scored within ±0%
  • rlr_06_status · scored within ±0%
  • rlr_07_status · scored within ±0%
  • rlr_08_status · scored within ±0%
  • rlr_09_status · scored within ±0%
  • readiness_code · scored within ±0%
  • required_findings_count · scored within ±0%
  • required_information_requests_count · scored within ±0%
  • required_carried_actions_count · scored within ±0%
  • design_density_mm_min · scored within ±2%
  • design_area_m2 · scored within ±2%
  • sprinkler_demand_l_min · scored within ±2%
  • hose_allowance_l_min · scored within ±2%
  • required_duration_min · scored within ±2%
  • required_volume_m3 · scored within ±2%
  • storage_volume_margin_m3 · scored within ±2%
  • pump_capacity_margin_l_min · scored within ±2%