LIVEdataset aec-bench@releasetasks 552models 18last submission · built
mechanicalwith-tool

Acoustic Review Repair Source Policy Package

Calculates source-bound acoustic review closure, source traceability, corrected margins, unresolved conflicts, and repair ledger completeness from a deterministic SSC-12 task-owned source pack.

with-tool: The model is given an executable Python calculator script.

How this task is generated

One template produces many comparable benchmark tasks while keeping the scoring contract fixed.

  1. 01

    Template

    The reusable contract shown on this page.

  2. 02

    Scenario

    An archetype and site context are sampled.

  3. 03

    Difficulty tier

    Inputs may be hidden at harder tiers.

  4. 04

    Task prompt

    The model responds with the declared outputs.

Parameters

Inputs the model receives, and the outputs it is scored on.

Inputs

14 inputs

Always given

Included directly in every task prompt.

14
Show 14 inputs
  • Closed review comments

    closed_review_comments

    COMMENT-12-REG-08 closed review comments

    6 count
  • Total review comments

    total_review_comments

    COMMENT-12-REG-08 total review comments

    6 count
  • Updated calculations

    updated_calculations

    RESPONSE-12-REPAIR-08 updated calculations

    5 count
  • Affected calculations

    affected_calculations

    COMMENT-12-REG-08 affected calculations

    5 count
  • Source referenced rows

    source_referenced_rows

    INDEX-12-SOURCE-08 referenced source rows

    7 count
  • Required source rows

    required_source_rows

    INDEX-12-SOURCE-08 required source rows

    7 count
  • Pre repair level

    pre_repair_level_dba

    SPEC-12-OCTAVE-08 pre-repair receiver level

    48.6 dBA
  • Post repair level

    post_repair_level_dba

    SPEC-12-OCTAVE-08 post-repair receiver level

    43.9 dBA
  • Noise criterion

    noise_criterion_dba

    CRIT-12-MATRIX-08 noise criterion

    45 dBA
  • Corrected vibration velocity

    corrected_vibration_velocity_mm_s

    RCV-12-PLAN-08 corrected vibration velocity

    0.28 mm/s
  • Vibration criterion

    vibration_criterion_mm_s

    CRIT-12-MATRIX-08 vibration criterion

    0.4 mm/s
  • Unresolved conflict

    unresolved_conflict_count

    RESPONSE-12-REPAIR-08 unresolved conflicts

    0 count
  • Complete repair ledger rows

    complete_repair_ledger_rows

    RESPONSE-12-REPAIR-08 complete repair ledger rows

    47 count
  • Required repair ledger rows

    required_repair_ledger_rows

    RESPONSE-12-REPAIR-08 required repair ledger rows

    50 count

Scored outputs

9 outputs

Review comment closure fraction

review_comment_closure_fraction

Closed review comments divided by total review comments

Scores if within ±0.1% of the reference value.

Affected calculation update fraction

affected_calculation_update_fraction

Updated calculations divided by affected calculations

Scores if within ±0.1% of the reference value.

Source traceability fraction

source_traceability_fraction

Referenced source rows divided by required source rows

Scores if within ±0.1% of the reference value.

Mitigation delta db

mitigation_delta_db

Pre-repair receiver level minus post-repair receiver level

Scores if within ±0.3% of the reference value.

Corrected noise margin db

corrected_noise_margin_db

Noise criterion minus post-repair receiver level

Scores if within ±0.3% of the reference value.

Vibration margin mm s

vibration_margin_mm_s

Vibration criterion minus corrected vibration velocity

Scores if within ±0.3% of the reference value.

Unresolved conflict count

unresolved_conflict_count

Unresolved source conflict count

Scores if within ±0.1% of the reference value.

Repair ledger completeness fraction

repair_ledger_completeness_fraction

Complete repair ledger rows divided by required rows

Scores if within ±0.1% of the reference value.

Overall pass score

overall_pass_score

1.0 when review, source, margin, conflict, and ledger checks pass

Scores if within ±0.1% of the reference value.

Difficulty

Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.

All inputs remain visible at every tier

For this template, difficulty scales through parameter and scenario ranges rather than hidden information.

easy:
All source-pack values given for the SSC-12 acoustic review repair package
medium:
All source-pack values given for the SSC-12 acoustic review repair package
hard:
All source-pack values given for the SSC-12 acoustic review repair package

Task bundle

The exact instruction and parameter contract used to generate this task, pinned to the published library source.

/workspace

  • instruction.md
  • acoustic-review-repair-source-policy-package_calc.py

Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.

1You are checking a task-owned synthetic SSC-12 acoustic review repair and source-policy package.2 3Use only the source pack values below for numeric grading. Acoustic review, source-index, comment-register, and criteria-matrix workflows provide context only; they are not extra data sources for this instance.4 5## Scene6 7- Product: `SSC-12-LH-08`8- Source index: `INDEX-12-SOURCE-08`9- Octave spectra: `SPEC-12-OCTAVE-08`10- Receiver plan: `RCV-12-PLAN-08`11- Comment register: `COMMENT-12-REG-08`12- Criteria matrix: `CRIT-12-MATRIX-08`13- Response memo: `RESPONSE-12-REPAIR-08`14 15Compute review comment closure, affected calculation updates, source traceability, mitigation delta, corrected noise margin, vibration margin, unresolved conflicts, repair ledger completeness, and pass score.16 17Write `/workspace/output.md` with a compact memo preserving the object IDs above. Include a source-boundary statement that this is a task-owned synthetic source pack.18 19Do not claim authority approval, accepted project evidence, software export validity, full standards compliance, executable real source-pack parsing, generated benchmark readiness, or benchmark readiness.20 21Include a fenced JSON block with exactly these numeric keys:22 23```json24{25 "review_comment_closure_fraction": <numeric_value>,26 "affected_calculation_update_fraction": <numeric_value>,27 "source_traceability_fraction": <numeric_value>,28 "mitigation_delta_db": <numeric_value>,29 "corrected_noise_margin_db": <numeric_value>,30 "vibration_margin_mm_s": <numeric_value>,31 "unresolved_conflict_count": <numeric_value>,32 "repair_ledger_completeness_fraction": <numeric_value>,33 "overall_pass_score": <numeric_value>34}35```36

Scenario archetypes

Each generated task is drawn from one of these realistic scenario bands.

Site contexts ground each scenario in a real locale the model can use to infer hidden values.

Ssc12 acoustic review case

ssc12_acoustic_review_case

Acoustic review repair and source-policy case

Parameter ranges
closed_review_comments
6
total_review_comments
6
updated_calculations
5
affected_calculations
5
source_referenced_rows
7
required_source_rows
7
pre_repair_level_dba
48.6
post_repair_level_dba
43.9
noise_criterion_dba
45
corrected_vibration_velocity_mm_s
0.28
vibration_criterion_mm_s
0.4
unresolved_conflict_count
0
complete_repair_ledger_rows
47
required_repair_ledger_rows
50

Example task

ssc12-acoustic-review-case-previewhard difficulty, all inputs given.

Acoustic review repair and source-policy case. Required outputs: review_comment_closure_fraction, affected_calculation_update_fraction, source_traceability_fraction, mitigation_delta_db, corrected_noise_margin_db, vibration_margin_mm_s

The model sees

Scenario context and visible inputs.

closed_review_comments
6 to 6 count
total_review_comments
6 to 6 count
updated_calculations
5 to 5 count
affected_calculations
5 to 5 count
source_referenced_rows
7 to 7 count
required_source_rows
7 to 7 count
pre_repair_level_dba
48.6 to 48.6 dBA
post_repair_level_dba
43.9 to 43.9 dBA
noise_criterion_dba
45 to 45 dBA
corrected_vibration_velocity_mm_s
0.28 to 0.28 mm/s
vibration_criterion_mm_s
0.4 to 0.4 mm/s
unresolved_conflict_count
0 to 0 count
complete_repair_ledger_rows
47 to 47 count
required_repair_ledger_rows
50 to 50 count

Executable tool: acoustic-review-repair-source-policy-package_calc.py

The model must infer

Inputs withheld at this difficulty.

Nothing. All inputs are supplied.

The model must produce

The scored JSON answer schema.

{
  "review_comment_closure_fraction": <number>,
  "affected_calculation_update_fraction": <number>,
  "source_traceability_fraction": <number>,
  "mitigation_delta_db": <number>,
  "corrected_noise_margin_db": <number>,
  "vibration_margin_mm_s": <number>,
  "unresolved_conflict_count": <number>,
  "repair_ledger_completeness_fraction": <number>,
  "overall_pass_score": <number>
}
  • review_comment_closure_fraction · scored within ±0.1%
  • affected_calculation_update_fraction · scored within ±0.1%
  • source_traceability_fraction · scored within ±0.1%
  • mitigation_delta_db · scored within ±0.3%
  • corrected_noise_margin_db · scored within ±0.3%
  • vibration_margin_mm_s · scored within ±0.3%
  • unresolved_conflict_count · scored within ±0.1%
  • repair_ledger_completeness_fraction · scored within ±0.1%
  • overall_pass_score · scored within ±0.1%