LIVEdataset aec-bench@releasetasks 552models 18last submission · built
civilwith-tool

Staging Review Response Negative Case Package

Calculates source-bound staging review closure, affected-check updates, source matching, revised environmental/traffic/power/tolerance margins, conflict count, and repair-ledger completeness from a deterministic SSC-16 task-owned source pack.

with-tool: The model is given an executable Python calculator script.

How this task is generated

One template produces many comparable benchmark tasks while keeping the scoring contract fixed.

  1. 01

    Template

    The reusable contract shown on this page.

  2. 02

    Scenario

    An archetype and site context are sampled.

  3. 03

    Difficulty tier

    Inputs may be hidden at harder tiers.

  4. 04

    Task prompt

    The model responds with the declared outputs.

Parameters

Inputs the model receives, and the outputs it is scored on.

Inputs

17 inputs

Always given

Included directly in every task prompt.

17
Show 17 inputs
  • Closed review comments

    closed_review_comments

    REVIEW-16-COMMENTS-08 closed review comments

    5 count
  • Total review comments

    total_review_comments

    REVIEW-16-COMMENTS-08 total review comments

    5 count
  • Updated affected checks

    updated_affected_checks

    CONTROL-16-SCHED-08 affected checks updated

    4 count
  • Total affected checks

    total_affected_checks

    CONTROL-16-SCHED-08 total affected checks

    4 count
  • Matching stage sources

    matching_stage_sources

    STAGE-16-REV-08 matching stage sources

    6 count
  • Referenced stage sources

    referenced_stage_sources

    STAGE-16-REV-08 referenced stage sources

    6 count
  • Revised basin capacity

    revised_basin_capacity_m3

    CONTROL-16-SCHED-08 revised basin capacity

    1335 m3
  • Required basin capacity

    required_basin_capacity_m3

    CRIT-16-MATRIX-08 required basin capacity

    1250 m3
  • Scheduled traffic device

    scheduled_traffic_device_count

    CONTROL-16-SCHED-08 scheduled traffic device count

    18 count
  • Inventoried traffic device

    inventoried_traffic_device_count

    DEVICE-16-INV-08 inventoried traffic device count

    18 count
  • Available temp power

    available_temp_power_w

    CONTROL-16-SCHED-08 available temporary power

    420 W
  • Revised temp load

    revised_temp_load_w

    CONTROL-16-SCHED-08 revised temporary load

    260 W
  • Allowed tolerance

    allowed_tolerance_mm

    CRIT-16-MATRIX-08 allowed construction tolerance

    15 mm
  • Observed tolerance

    observed_tolerance_mm

    DEVICE-16-INV-08 observed construction tolerance

    9 mm
  • Unresolved conflict

    unresolved_conflict_count

    RESPONSE-16-REPAIR-08 unresolved conflict count

    0 count
  • Completed repair fields

    completed_repair_fields

    RESPONSE-16-REPAIR-08 completed repair-ledger fields

    23 count
  • Required repair fields

    required_repair_fields

    RESPONSE-16-REPAIR-08 required repair-ledger fields

    25 count

Scored outputs

10 outputs

Review comment closure fraction

review_comment_closure_fraction

Closed review comments divided by total review comments

Scores if within ±0.1% of the reference value.

Affected check update fraction

affected_check_update_fraction

Updated affected checks divided by total affected checks

Scores if within ±0.1% of the reference value.

Stage source match fraction

stage_source_match_fraction

Matching stage sources divided by referenced stage sources

Scores if within ±0.1% of the reference value.

Sediment basin margin m3

sediment_basin_margin_m3

Revised basin capacity minus required basin capacity

Scores if within ±3% of the reference value.

Traffic device delta count

traffic_device_delta_count

Absolute difference between scheduled and inventoried traffic devices

Scores if within ±0.1% of the reference value.

Power headroom w

power_headroom_w

Available temporary power minus revised load

Scores if within ±3% of the reference value.

Tolerance margin

tolerance_margin_mm

Allowed tolerance minus observed tolerance

Scores if within ±3% of the reference value.

Unresolved conflict count

unresolved_conflict_count

Unresolved source conflict count

Scores if within ±0.1% of the reference value.

Repair ledger completeness fraction

repair_ledger_completeness_fraction

Completed repair-ledger fields divided by required fields

Scores if within ±0.1% of the reference value.

Overall pass score

overall_pass_score

1.0 when review, source, margin, conflict, and repair checks pass

Scores if within ±0.1% of the reference value.

Difficulty

Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.

All inputs remain visible at every tier

For this template, difficulty scales through parameter and scenario ranges rather than hidden information.

easy:
All source-pack values given for the SSC-16 staging review response package
medium:
All source-pack values given for the SSC-16 staging review response package
hard:
All source-pack values given for the SSC-16 staging review response package

Task bundle

The exact instruction and parameter contract used to generate this task, pinned to the published library source.

/workspace

  • instruction.md
  • staging-review-response-negative-case-package_calc.py

Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.

1You are a construction staging reviewer checking a task-owned synthetic SSC-16 package for review closeout, affected calculation updates, stage-source identity, and repair-ledger completeness.2 3Use only the task-owned synthetic source pack values shown below for numeric grading. External review, hold-point, inspection, and staging workflows shape the context only; they are not extra data sources for this instance.4 5## Source Objects6 7- Product: `SSC-16-LH-08`8- Review comments: `REVIEW-16-COMMENTS-08`9- Revised staging plan: `STAGE-16-REV-08`10- Control schedule: `CONTROL-16-SCHED-08`11- Device inventory: `DEVICE-16-INV-08`12- Criteria matrix: `CRIT-16-MATRIX-08`13- Repair response memo: `RESPONSE-16-REPAIR-08`14 15## Source Values16 17| Item | Value |18|------|-------|19| Closed review comments | {{ closed_review_comments }} |20| Total review comments | {{ total_review_comments }} |21| Updated affected checks | {{ updated_affected_checks }} |22| Total affected checks | {{ total_affected_checks }} |23| Matching stage sources | {{ matching_stage_sources }} |24| Referenced stage sources | {{ referenced_stage_sources }} |25| Revised basin capacity | {{ revised_basin_capacity_m3 }} m3 |26| Required basin capacity | {{ required_basin_capacity_m3 }} m3 |27| Scheduled traffic devices | {{ scheduled_traffic_device_count }} |28| Inventoried traffic devices | {{ inventoried_traffic_device_count }} |29| Available temporary power | {{ available_temp_power_w }} W |30| Revised temporary load | {{ revised_temp_load_w }} W |31| Allowed tolerance | {{ allowed_tolerance_mm }} mm |32| Observed tolerance | {{ observed_tolerance_mm }} mm |33| Unresolved conflicts | {{ unresolved_conflict_count }} |34| Completed repair-ledger fields | {{ completed_repair_fields }} |35| Required repair-ledger fields | {{ required_repair_fields }} |36 37## Required Checks38 39- Review comments and affected calculations must be fully closed or updated.40- Stage source matching confirms the response uses the same source objects as the revised stage.41- Environmental, traffic, power, and tolerance margins must remain non-negative after repair.42- The repair ledger must have no unresolved conflicts and at least 0.9 completeness.43 44## Output Format45 46Write a compact memo to `/workspace/output.md`. Include a source-boundary statement that this is a task-owned synthetic source pack, preserve the source object IDs above, and state whether the baseline checks pass.47 48Do not claim authority approval, accepted project evidence, software export validity, full standards compliance, executable source-pack hardening, generated benchmark readiness, or benchmark readiness.49 50Include a fenced JSON block with exactly these numeric keys:51 52```json53{54 "review_comment_closure_fraction": <numeric_value>,55 "affected_check_update_fraction": <numeric_value>,56 "stage_source_match_fraction": <numeric_value>,57 "sediment_basin_margin_m3": <numeric_value>,58 "traffic_device_delta_count": <numeric_value>,59 "power_headroom_w": <numeric_value>,60 "tolerance_margin_mm": <numeric_value>,61 "unresolved_conflict_count": <numeric_value>,62 "repair_ledger_completeness_fraction": <numeric_value>,63 "overall_pass_score": <numeric_value>64}65```66

Scenario archetypes

Each generated task is drawn from one of these realistic scenario bands.

Site contexts ground each scenario in a real locale the model can use to infer hidden values.

Ssc16 staging review case

ssc16_staging_review_case

Staging review response and negative-case repair case

Parameter ranges
closed_review_comments
5
total_review_comments
5
updated_affected_checks
4
total_affected_checks
4
matching_stage_sources
6
referenced_stage_sources
6
revised_basin_capacity_m3
1335
required_basin_capacity_m3
1250
scheduled_traffic_device_count
18
inventoried_traffic_device_count
18
available_temp_power_w
420
revised_temp_load_w
260
allowed_tolerance_mm
15
observed_tolerance_mm
9
unresolved_conflict_count
0
completed_repair_fields
23
required_repair_fields
25

Example task

ssc16-staging-review-case-previewhard difficulty, all inputs given.

Staging review response and negative-case repair case. Required outputs: review_comment_closure_fraction, affected_check_update_fraction, stage_source_match_fraction, sediment_basin_margin_m3, traffic_device_delta_count, power_headroom_w

The model sees

Scenario context and visible inputs.

closed_review_comments
5 to 5 count
total_review_comments
5 to 5 count
updated_affected_checks
4 to 4 count
total_affected_checks
4 to 4 count
matching_stage_sources
6 to 6 count
referenced_stage_sources
6 to 6 count
revised_basin_capacity_m3
1335 to 1335 m3
required_basin_capacity_m3
1250 to 1250 m3
scheduled_traffic_device_count
18 to 18 count
inventoried_traffic_device_count
18 to 18 count
available_temp_power_w
420 to 420 W
revised_temp_load_w
260 to 260 W
allowed_tolerance_mm
15 to 15 mm
observed_tolerance_mm
9 to 9 mm
unresolved_conflict_count
0 to 0 count
completed_repair_fields
23 to 23 count
required_repair_fields
25 to 25 count

Executable tool: staging-review-response-negative-case-package_calc.py

The model must infer

Inputs withheld at this difficulty.

Nothing. All inputs are supplied.

The model must produce

The scored JSON answer schema.

{
  "review_comment_closure_fraction": <number>,
  "affected_check_update_fraction": <number>,
  "stage_source_match_fraction": <number>,
  "sediment_basin_margin_m3": <number>,
  "traffic_device_delta_count": <number>,
  "power_headroom_w": <number>,
  "tolerance_margin_mm": <number>,
  "unresolved_conflict_count": <number>,
  "repair_ledger_completeness_fraction": <number>,
  "overall_pass_score": <number>
}
  • review_comment_closure_fraction · scored within ±0.1%
  • affected_check_update_fraction · scored within ±0.1%
  • stage_source_match_fraction · scored within ±0.1%
  • sediment_basin_margin_m3 · scored within ±3%
  • traffic_device_delta_count · scored within ±0.1%
  • power_headroom_w · scored within ±3%
  • tolerance_margin_mm · scored within ±3%
  • unresolved_conflict_count · scored within ±0.1%
  • repair_ledger_completeness_fraction · scored within ±0.1%
  • overall_pass_score · scored within ±0.1%