LIVEdataset aec-bench@releasetasks 552models 18last submission · built
electricalwith-tool

Visual Systems Review Repair Package

Calculates source-bound visual systems review metrics from a deterministic SSC-13 task-owned source pack. The template combines review comments, revised layout, device schedule, calculation trace, criteria matrix, and repair response.

with-tool: The model is given an executable Python calculator script.

How this task is generated

One template produces many comparable benchmark tasks while keeping the scoring contract fixed.

  1. 01

    Template

    The reusable contract shown on this page.

  2. 02

    Scenario

    An archetype and site context are sampled.

  3. 03

    Difficulty tier

    Inputs may be hidden at harder tiers.

  4. 04

    Task prompt

    The model responds with the declared outputs.

Parameters

Inputs the model receives, and the outputs it is scored on.

Inputs

16 inputs

Always given

Included directly in every task prompt.

16
Show 16 inputs
  • Closed review comments

    closed_review_comments

    REVIEW-13-COMMENTS-08 closed review comments

    5 count
  • Total review comments

    total_review_comments

    REVIEW-13-COMMENTS-08 total review comments

    5 count
  • Updated affected checks

    updated_affected_checks

    CALC-13-TRACE-08 updated affected checks

    4 count
  • Required affected checks

    required_affected_checks

    CRIT-13-MATRIX-08 required affected checks

    4 count
  • Revised minimum

    revised_minimum_lux

    LAYOUT-13-REV-08 revised minimum lighting

    16.1 lux
  • Required minimum

    required_minimum_lux

    CRIT-13-MATRIX-08 required minimum lighting

    15 lux
  • Cctv horizontal pixels

    cctv_horizontal_pixels

    DEVICE-13-SCHED-08 revised CCTV horizontal pixels

    1920 px
  • Revised target width

    revised_target_width_m

    LAYOUT-13-REV-08 revised CCTV target width

    24 m
  • Required ppm

    required_ppm

    CRIT-13-MATRIX-08 required CCTV PPM

    70
  • Revised network load

    revised_network_load_mbps

    DEVICE-13-SCHED-08 revised network load

    38 Mbps
  • Network capacity

    network_capacity_mbps

    DEVICE-13-SCHED-08 network capacity

    50 Mbps
  • Revised poe load

    revised_poe_load_w

    DEVICE-13-SCHED-08 revised PoE load

    92 W
  • Poe budget

    poe_budget_w

    DEVICE-13-SCHED-08 PoE budget

    120 W
  • Unresolved conflict

    unresolved_conflict_count

    RESPONSE-13-REPAIR-08 unresolved conflicts

    0 count
  • Completed repair memo sections

    completed_repair_memo_sections

    RESPONSE-13-REPAIR-08 completed memo sections

    9 count
  • Required repair memo sections

    required_repair_memo_sections

    RESPONSE-13-REPAIR-08 required memo sections

    10 count

Scored outputs

10 outputs

Review comment closure fraction

review_comment_closure_fraction

Closed review comments divided by total comments

Scores if within ±0.3% of the reference value.

Affected check update fraction

affected_check_update_fraction

Updated affected checks divided by required affected checks

Scores if within ±0.3% of the reference value.

Lighting minimum margin lux

lighting_minimum_margin_lux

Revised minimum lighting minus required minimum lighting

Scores if within ±0.3% of the reference value.

Revised cctv pixels per

revised_cctv_pixels_per_m

Revised CCTV pixels per metre

Scores if within ±0.3% of the reference value.

Cctv ppm margin

cctv_ppm_margin

Revised CCTV PPM minus required PPM

Scores if within ±0.3% of the reference value.

Network headroom mbps

network_headroom_mbps

Network capacity minus revised network load

Scores if within ±0.3% of the reference value.

Poe headroom w

poe_headroom_w

PoE budget minus revised PoE load

Scores if within ±0.3% of the reference value.

Unresolved conflict count

unresolved_conflict_count

Unresolved source conflict count

Scores if within ±0.1% of the reference value.

Repair memo completeness fraction

repair_memo_completeness_fraction

Completed repair memo sections divided by required sections

Scores if within ±0.3% of the reference value.

Overall pass score

overall_pass_score

1.0 when review repair checks pass

Scores if within ±0.1% of the reference value.

Difficulty

Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.

All inputs remain visible at every tier

For this template, difficulty scales through parameter and scenario ranges rather than hidden information.

easy:
All source-pack values given for the SSC-13 visual systems repair package
medium:
All source-pack values given for the SSC-13 visual systems repair package
hard:
All source-pack values given for the SSC-13 visual systems repair package

Task bundle

The exact instruction and parameter contract used to generate this task, pinned to the published library source.

/workspace

  • instruction.md
  • visual-systems-review-repair-package_calc.py

Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.

1You are an electrical visual systems reviewer checking a task-owned synthetic SSC-13 review and repair package.2 3Use only the task-owned synthetic source pack values shown below for numeric grading. External review-management, lighting, CCTV, ITS, and network tools shape the workflow context only; they are not extra data sources for this instance.4 5## Scene6 7- Product: `SSC-13-LH-08`8- Review comments: `REVIEW-13-COMMENTS-08`9- Revised layout: `LAYOUT-13-REV-08`10- Revised device schedule: `DEVICE-13-SCHED-08`11- Calculation trace: `CALC-13-TRACE-08`12- Criteria matrix: `CRIT-13-MATRIX-08`13- Repair response: `RESPONSE-13-REPAIR-08`14 15All checks use the same comment register, revised layout, device schedule, affected calculation trace, criteria matrix, and response ledger.16 17## Source Values18 19| Item | Value |20|------|-------|21| Closed review comments | {{ closed_review_comments }} of {{ total_review_comments }} |22| Updated affected checks | {{ updated_affected_checks }} of {{ required_affected_checks }} |23| Revised minimum lighting | {{ revised_minimum_lux }} lux |24| Required minimum lighting | {{ required_minimum_lux }} lux |25| CCTV pixels and revised target width | {{ cctv_horizontal_pixels }} px / {{ revised_target_width_m }} m |26| Required PPM | {{ required_ppm }} |27| Revised network load and capacity | {{ revised_network_load_mbps }} Mbps / {{ network_capacity_mbps }} Mbps |28| Revised PoE load and budget | {{ revised_poe_load_w }} W / {{ poe_budget_w }} W |29| Unresolved conflicts | {{ unresolved_conflict_count }} |30| Repair memo sections | {{ completed_repair_memo_sections }} of {{ required_repair_memo_sections }} |31 32## Output Format33 34Write a compact visual systems review response to `/workspace/output.md`. Include a source-boundary statement that this is a task-owned synthetic source pack. Do not claim authority approval, accepted project evidence, software export validity, full standards compliance, executable real source-pack parsing, generated benchmark readiness, or benchmark readiness.35 36Include a fenced JSON block with exactly these numeric keys:37 38```json39{40 "review_comment_closure_fraction": <numeric_value>,41 "affected_check_update_fraction": <numeric_value>,42 "lighting_minimum_margin_lux": <numeric_value>,43 "revised_cctv_pixels_per_m": <numeric_value>,44 "cctv_ppm_margin": <numeric_value>,45 "network_headroom_mbps": <numeric_value>,46 "poe_headroom_w": <numeric_value>,47 "unresolved_conflict_count": <numeric_value>,48 "repair_memo_completeness_fraction": <numeric_value>,49 "overall_pass_score": <numeric_value>50}51```52

Scenario archetypes

Each generated task is drawn from one of these realistic scenario bands.

Site contexts ground each scenario in a real locale the model can use to infer hidden values.

Ssc13 visual repair case

ssc13_visual_repair_case

Visual systems review and repair case

ssc13-visual-repair
Parameter ranges

Example task

ssc13-visual-repair-ssc13-visual-repair-case-previewhard difficulty, all inputs given.

Visual systems review and repair case. ssc13-visual-repair. Required outputs: review_comment_closure_fraction, affected_check_update_fraction, lighting_minimum_margin_lux, revised_cctv_pixels_per_m, cctv_ppm_margin, network_headroom_mbps

The model sees

Scenario context and visible inputs.

closed_review_comments
5 to 5 count
total_review_comments
5 to 5 count
updated_affected_checks
4 to 4 count
required_affected_checks
4 to 4 count
revised_minimum_lux
16.1 to 16.1 lux
required_minimum_lux
15 to 15 lux
cctv_horizontal_pixels
1920 to 1920 px
revised_target_width_m
24 to 24 m
required_ppm
70 to 70
revised_network_load_mbps
38 to 38 Mbps
network_capacity_mbps
50 to 50 Mbps
revised_poe_load_w
92 to 92 W
poe_budget_w
120 to 120 W
unresolved_conflict_count
0 to 0 count
completed_repair_memo_sections
9 to 9 count
required_repair_memo_sections
10 to 10 count

Executable tool: visual-systems-review-repair-package_calc.py

The model must infer

Inputs withheld at this difficulty.

Nothing. All inputs are supplied.

The model must produce

The scored JSON answer schema.

{
  "review_comment_closure_fraction": <number>,
  "affected_check_update_fraction": <number>,
  "lighting_minimum_margin_lux": <number>,
  "revised_cctv_pixels_per_m": <number>,
  "cctv_ppm_margin": <number>,
  "network_headroom_mbps": <number>,
  "poe_headroom_w": <number>,
  "unresolved_conflict_count": <number>,
  "repair_memo_completeness_fraction": <number>,
  "overall_pass_score": <number>
}
  • review_comment_closure_fraction · scored within ±0.3%
  • affected_check_update_fraction · scored within ±0.3%
  • lighting_minimum_margin_lux · scored within ±0.3%
  • revised_cctv_pixels_per_m · scored within ±0.3%
  • cctv_ppm_margin · scored within ±0.3%
  • network_headroom_mbps · scored within ±0.3%
  • poe_headroom_w · scored within ±0.3%
  • unresolved_conflict_count · scored within ±0.1%
  • repair_memo_completeness_fraction · scored within ±0.3%
  • overall_pass_score · scored within ±0.1%