Datum trace score
datum_trace_score
Datum items traced divided by required datum items
Scores if within ±0.3% of the reference value.
Calculates source-bound marine asset source-policy review metrics from a deterministic SSC-04 task-owned source pack. The template combines datum trace, criteria resolution, asset schedule matching, calculation trace, comment resolution, authority partition, unsupported-source count, response completeness, evidence boundary score, and an overall synthetic pass score.
with-tool: The model is given an executable Python calculator script.
Standards
One template produces many comparable benchmark tasks while keeping the scoring contract fixed.
01
The reusable contract shown on this page.
02
An archetype and site context are sampled.
03
Inputs may be hidden at harder tiers.
04
The model responds with the declared outputs.
Inputs the model receives, and the outputs it is scored on.
15 inputs
Included directly in every task prompt.
Datum items traced
datum_items_traced
DATUM-04-STATEMENT-08 traced datum items
Required datum items
required_datum_items
DATUM-04-STATEMENT-08 required datum items
Resolved criteria items
resolved_criteria_items
CRITERIA-04-MATRIX-08 resolved criteria items
Criteria items
criteria_items
CRITERIA-04-MATRIX-08 criteria items
Matching asset rows
matching_asset_rows
ASSET-04-MARINE-08 matching asset rows
Asset schedule rows
asset_schedule_rows
ASSET-04-MARINE-08 asset schedule rows
Traced calculation rows
traced_calculation_rows
CALC-04-APPENDIX-08 traced calculation rows
Calculation rows
calculation_rows
CALC-04-APPENDIX-08 calculation rows
Resolved comments
resolved_comments
COMMENT-04-REGISTER-08 resolved comments
Review comments
review_comments
COMMENT-04-REGISTER-08 review comments
Separated authority roles
separated_authority_roles
MEMO-04-SOURCE-08 separated authority roles
Required authority roles
required_authority_roles
MEMO-04-SOURCE-08 required authority roles
Unsupported source value count
unsupported_source_value_count
MEMO-04-SOURCE-08 unsupported source value count
Response sections
response_sections
MEMO-04-SOURCE-08 response sections
Required response sections
required_response_sections
MEMO-04-SOURCE-08 required response sections
10 outputs
datum_trace_score
Datum items traced divided by required datum items
Scores if within ±0.3% of the reference value.
criteria_resolution_fraction
Resolved criteria items divided by criteria items
Scores if within ±0.3% of the reference value.
asset_schedule_match_fraction
Matching asset rows divided by asset schedule rows
Scores if within ±0.3% of the reference value.
calculation_trace_fraction
Traced calculation rows divided by calculation rows
Scores if within ±0.3% of the reference value.
comment_resolution_fraction
Resolved comments divided by review comments
Scores if within ±0.3% of the reference value.
authority_partition_score
Separated authority roles divided by required authority roles
Scores if within ±0.3% of the reference value.
unsupported_source_value_count
Unsupported source value count
Scores if within ±0.1% of the reference value.
response_completeness_score
Response sections divided by required response sections
Scores if within ±0.3% of the reference value.
evidence_boundary_score
Average evidence-boundary score
Scores if within ±0.3% of the reference value.
overall_pass_score
1.0 when source-policy review checks pass
Scores if within ±0.1% of the reference value.
Each template is sampled at three tiers. Harder tiers may hide inputs, forcing the model to infer them from the scenario description.
For this template, difficulty scales through parameter and scenario ranges rather than hidden information.
The exact instruction and parameter contract used to generate this task, pinned to the published library source.
/workspace
Teal lines show Jinja input conditions, not task visibility policy. A line renders only when that input or tool is visible.
1You are a civil/marine reviewer checking a task-owned synthetic SSC-04 marine asset source-policy and review packet.2 3Use only the task-owned synthetic source pack values shown below for numeric grading. External datum, criteria, asset-register, calculation-appendix, comment-register, and authority-boundary workflows shape the practice context only; they are not extra data sources for this instance.4 5## Scene6 7- Product family: `SSC-04-LH-08`8- Datum statement: `DATUM-04-STATEMENT-08`9- Criteria matrix: `CRITERIA-04-MATRIX-08`10- Marine asset schedule: `ASSET-04-MARINE-08`11- Calculation appendix: `CALC-04-APPENDIX-08`12- Comment register: `COMMENT-04-REGISTER-08`13- Source-policy memo: `MEMO-04-SOURCE-08`14 15## Source Values16 17| Item | Value |18|------|-------|19| Datum items traced | {{ datum_items_traced }} |20| Required datum items | {{ required_datum_items }} |21| Resolved criteria items | {{ resolved_criteria_items }} |22| Criteria items | {{ criteria_items }} |23| Matching asset rows | {{ matching_asset_rows }} |24| Asset schedule rows | {{ asset_schedule_rows }} |25| Traced calculation rows | {{ traced_calculation_rows }} |26| Calculation rows | {{ calculation_rows }} |27| Resolved comments | {{ resolved_comments }} |28| Review comments | {{ review_comments }} |29| Separated authority roles | {{ separated_authority_roles }} |30| Required authority roles | {{ required_authority_roles }} |31| Unsupported source value count | {{ unsupported_source_value_count }} |32| Response sections | {{ response_sections }} |33| Required response sections | {{ required_response_sections }} |34 35## Checks36 37- Datum trace score equals datum items traced divided by required datum items.38- Criteria resolution fraction equals resolved criteria items divided by criteria items.39- Asset schedule match fraction equals matching asset rows divided by asset schedule rows.40- Calculation trace fraction equals traced calculation rows divided by calculation rows.41- Comment resolution fraction equals resolved comments divided by review comments.42- Authority partition score equals separated authority roles divided by required authority roles.43- Response completeness score equals response sections divided by required response sections.44- Evidence boundary score is the average of datum trace, criteria resolution, asset match, calculation trace, comment resolution, authority partition, and response completeness scores.45- Overall pass score is `1.0` only when datum, criteria, and authority partition checks are complete, unsupported source values are zero, and response completeness is at least 0.9; otherwise it is `0.0`.46 47## Output Format48 49Write a compact source-policy review memo to `/workspace/output.md`. Include a source-boundary statement that this is a task-owned synthetic source pack. Preserve the object IDs above and state whether the baseline source pack passes the current docs-only checks.50 51Do not claim authority approval, accepted project evidence, software export validity, full standards compliance, executable real source-pack parsing, generated benchmark readiness, or benchmark readiness.52 53Include a fenced JSON block with exactly these numeric keys:54 55```json56{57 "datum_trace_score": <numeric_value>,58 "criteria_resolution_fraction": <numeric_value>,59 "asset_schedule_match_fraction": <numeric_value>,60 "calculation_trace_fraction": <numeric_value>,61 "comment_resolution_fraction": <numeric_value>,62 "authority_partition_score": <numeric_value>,63 "unsupported_source_value_count": <numeric_value>,64 "response_completeness_score": <numeric_value>,65 "evidence_boundary_score": <numeric_value>,66 "overall_pass_score": <numeric_value>67}68```69 Each generated task is drawn from one of these realistic scenario bands.
Site contexts ground each scenario in a real locale the model can use to infer hidden values.
ssc04_marine_source_policy_case
Marine asset source-policy and review packet case
ssc04-marine-source-policy-case-preview — hard difficulty, all inputs given.
Marine asset source-policy and review packet case. Required outputs: datum_trace_score, criteria_resolution_fraction, asset_schedule_match_fraction, calculation_trace_fraction, comment_resolution_fraction, authority_partition_score
Scenario context and visible inputs.
Executable tool: marine-asset-source-policy-review-package_calc.py
Inputs withheld at this difficulty.
Nothing. All inputs are supplied.
The scored JSON answer schema.
{
"datum_trace_score": <number>,
"criteria_resolution_fraction": <number>,
"asset_schedule_match_fraction": <number>,
"calculation_trace_fraction": <number>,
"comment_resolution_fraction": <number>,
"authority_partition_score": <number>,
"unsupported_source_value_count": <number>,
"response_completeness_score": <number>,
"evidence_boundary_score": <number>,
"overall_pass_score": <number>
}