OASIS OASIS
  • Home
  • Explore
    • Tour overview
    • CLI & TUI
    • Candle CLI & MCP
    • Elephant ingestion
    • MAPLES
    • Improve a rubric
  • Technical Report
  • PDF
  • Demo repository

Improve a rubric, one decision at a time

Watch v1 become v3 through prepared feedback, human decisions, and a dry run

Follow a prepared bicycle brake-adjustment rubric from its vague first version through Rubric Maker-style feedback, selected revisions, trial feedback, and a v3 candidate ready for expert review.

This guide uses a newly written bicycle brake-adjustment rubric to make the method easy to see without clinical, learner, or other real-person data. Follow the version strip to see exactly how v1 becomes v2, how a dry run exposes another problem, and how that feedback produces v3. The content is prepared and illustrative, mirroring the released agent-skill workflow. Using the page does not start an agent or model, run a grader, or contact MAPLES or OASIS.

Scope note: Rubric Maker v0.1.0 is an installable agent plugin focused on OSCE rubrics. Bicycle maintenance is outside the plugin’s documented OSCE scope. It is used here only to teach the general review loop; a qualified domain expert still needs to approve any rubric before use.

Interactive rubric guide

Start with the rubric you actually have

Preserve the original wording as version 1 so every later change is visible and reversible.

Prepared example · no Rubric Maker run

Rubric evolution

Full path: v1 baseline → prepared review feedback → human-selected v2 → prepared dry-run feedback → v3 candidate.

  1. Stage 1 v1 Baseline Original wording preserved
  2. Stage 2 Round 1 Review feedback Prepared · not a skill run
  3. Stage 3 v2 Selected anchors Human decision · not saved
  4. Stage 4 Round 2 Dry-run feedback Prepared · not a grading run
  5. Stage 5 v3 Candidate Expert approval still required
Step 1 · Baseline

Import faithfully before improving

In a real agent session, rubric-import can bring existing material into Rubric Maker YAML or JSON while preserving its wording and structure. For this walkthrough, the version below was written specifically for the page.

Prepared source artifact · bicycle-brake-rubric-v1.yaml

Mechanical rim-brake adjustmentv1 · baseline
Prepare safely
1 · unsafe   2 · partly safe   3 · safe
Adjust the brake correctly
1 · incorrect   2 · mostly correct   3 · correct
Test the brake
1 · not tested   2 · partly tested   3 · tested

The displayed values map to the plugin schema’s contiguous Score1, Score2, and Score3 keys.

Why keep v1? A faithful baseline prevents a helpful-looking rewrite from silently changing the task, score scale, or evidence you intended to assess.
Step 2 · Focus

Source: prepared review framing, authored with Codex assistance; no Rubric Maker skill was run.

Name the decision a rater cannot make reliably

“Adjust the brake correctly” sounds reasonable, but it does not say what a reviewer should see or how “mostly correct” differs from “correct.” We will improve that one decision first.

Review question What observable evidence separates a usable adjustment from one that still needs work?

Evidence the task can actually reveal

  • Each pad contacts the rim, not the tire.
  • Whether both pads clear the rim after the lever is released.
  • Where contact recurs as the wheel turns by hand.
  • The lever applies the brake before reaching the handlebar.

These details are authored for the example. They are not advice for repairing a bicycle and still require expert review.

Step 3 · Review

Make the critique inspectable

Rubric Maker’s osce-rubric-review skill examines alignment, safety, observability, objectivity, feasibility, and reliability. The notes below are prepared examples of that structure—not output from a Rubric Maker review run.

Feedback source: prepared Rubric Maker-style review notes, newly authored with Codex assistance; illustrative, not tool output.

AlignmentKeep

The item matches the stated brake-adjustment task.

SafetyImprove

No stop rule covers a damaged rim or frayed cable.

ObservabilityImprove

“Correctly” does not identify visible evidence.

ObjectivityImprove

“Mostly” leaves too much room for interpretation.

FeasibilityKeep

The core checks can be observed without special equipment.

ReliabilityImprove

Two reviewers could score the same adjustment differently.

Priority: define the observable pass condition and add a safety stop before polishing style.
Step 4 · Compare

Review suggestions one change at a time

The released workflow can return one structured suggestion for each field or score anchor being changed. These four suggestions and reasons were prewritten for this guide; RM-01 through RM-04 are editorial walkthrough labels, not required fields in the v0.1.0 suggestion contract.

Feedback source: prepared Rubric Maker-style suggestions, authored with Codex assistance; each proposal targets one explicit artifact location.

RM-01 · Score1 safety boundary
1:ScoringLogic.Score1
Recommended
Current

Score 1: incorrect

Proposed

Score 1 and stop when the observation shows a damaged rim, a frayed cable, tire contact, or a brake that does not engage.

Reason: makes the lowest anchor and its safety boundary explicit.

RM-02 · Score2 middle anchor
1:ScoringLogic.Score2
Recommended
Current

Score 2: mostly correct

Proposed

Score 2 when no stop condition is present but at least one passing check is missing.

Reason: gives the middle anchor a concrete distinction from both stopping and passing.

RM-03 · Score3 passing anchor
1:ScoringLogic.Score3
Recommended
Current

Score 3: correct

Proposed

Score 3 when both pads contact the rim below the tire, release fully, the wheel turns without continuous scraping, and the lever does not reach the handlebar.

Reason: ties the passing score to evidence that two reviewers can inspect.

RM-04 · Technique measurement
1:Technique
Question
Current

Technique is empty in the prepared source.

Proposed

Require exactly 1.5 mm of clearance at both pads.

Reason to pause: the proposed precision needs a measurement method the task and observation setup do not provide.

Step 5 · Decide

Accept the evidence, not the confidence

In a real agent conversation, the reviewer can name the suggestions or rubric locations to apply. For this prepared path, the human decision is fixed and visible: apply RM-01, RM-02, and RM-03; leave RM-04 out. The buttons here are walkthrough controls only; they do not imitate native Rubric Maker, Wayfinder, MAPLES, or SimRubrics controls. Use them to compare the preserved v1 fields with the resulting v2 candidate. Nothing is saved or validated.

Decision source: a human applies or declines the prepared suggestions; osce-rubric-transform would apply only those selected changes in a real agent workflow.

  • Apply RM-01 · 1:ScoringLogic.Score1
  • Apply RM-02 · 1:ScoringLogic.Score2
  • Apply RM-03 · 1:ScoringLogic.Score3
  • Leave out RM-04 · 1:Technique

Walkthrough candidate v2: RM-01, RM-02, and RM-03 applied; RM-04 left out. Nothing is saved or validated.

Mechanical rim-brake adjustmentv2 · walkthrough candidate

3 prepared changes applied · not saved or validated

Prepare safelyUnchanged from v1

Focus criterion · Adjust the brake correctly

Score1Changed in v2
Stop when the observation shows a damaged rim, a frayed cable, tire contact, or a brake that does not engage.Incorrect.
Score2Changed in v2
No stop condition is present, but at least one passing check is missing.Mostly correct.
Score3Changed in v2
Both pads contact the rim below the tire, release fully, the wheel turns without continuous scraping, and the lever does not reach the handlebar.Correct.
TechniquePreserved from v1
Require exactly 1.5 mm of clearance at both pads.Empty; the prepared precision proposal is left out.

Test the brakeUnchanged from v1

Evolution checkpoint: the default prepared decision changes three separate score-anchor fields and preserves Technique. The original bicycle-brake-rubric-v1.yaml remains untouched.
Dry-run evidence rule: when the evidence is insufficient to score, mark the grade-sheet item unscorable: true, state why, and route it for review rather than guessing. This is not a fourth rubric score anchor.

In a real workflow, ask osce-rubric-transform to write bicycle-brake-rubric-v2.yaml, then validate it with Rubric Maker’s bundled support. This page does neither operation.

Step 6 · Trial

Look for grading friction, not a flattering score

Rubric Maker v0.1.0 includes grading-dry-run and evaluate-dry-run. The observations and scores below are editorial illustrations, not outputs from a Rubric Maker skill run or a recorded grading run.

Feedback source: prepared grading-dry-run-style results and evaluate-dry-run-style review, authored with Codex assistance; no grader or model ran.

Prepared trial of the revised brake-adjustment item
Synthetic observation Prepared score What the anchor reveals
T-01 · Pads meet the rim below the tire; both release; no continuous scrape; lever stops short of the handlebar. 3 · meets All stated evidence is present.
T-02 · Left pad touches the tire when the lever is applied. 1 · stop The safety boundary is visible.
T-03 · Pads meet the rim and the lever holds, but one pad remains in contact after release. 2 · revise The middle score now has a concrete defect.
T-04 · The right pad brushes once per wheel revolution, but the view does not show the pad gap after release. unscorable: true “Continuous scraping” invites a 2-versus-3 disagreement, while the available view cannot resolve pad clearance.
Prepared feedback F1 · 1:ScoringLogic.Score3Replace “release fully” and “the wheel turns without continuous scraping” with one directly observable release-and-clearance clause.

T-04 exposes both a wording disagreement and an evidence gap. It remains unscorable rather than being forced into Score2 or Score3.

Step 7 · Revise

Make the smallest defensible change

The prepared dry-run feedback did not justify a wholesale rewrite. A human accepts F1 for 1:ScoringLogic.Score3, then a real workflow could use osce-rubric-transform to write a new artifact. This page shows the prepared result; it does not run that skill.

Refinement source: human acceptance of prepared feedback F1; the v3 wording below was authored with Codex assistance and is not saved or validated.

v2 · Score3

“Both pads contact the rim below the tire, release fully, the wheel turns without continuous scraping, and the lever does not reach the handlebar.”

v3 · Score3

“Both pads contact the rim below the tire; after release, a visible gap returns at both pads and remains for one hand turn; the lever does not reach the handlebar.”

Mechanical rim-brake adjustmentv3 · prepared candidate

Prepare safelyUnchanged from v1

Focus criterion · Adjust the brake correctly

Score1Carried from v2
Stop for a damaged rim, frayed cable, tire contact, or a brake that does not engage.
Score2Carried from v2
No stop condition is present, but at least one passing check is missing.
Score3Changed in v3 from F1
Both pads contact the rim below the tire; after release, a visible gap returns at both pads and remains for one hand turn; the lever does not reach the handlebar.
TechniquePreserved
Empty; no unsupported precision was added.

Test the brakeUnchanged from v1

Prepared T-04 rerun: the revised anchor now asks for a visible gap through one hand turn. The old observation still lacks that view, so it correctly remains unscorable: true; a new observation with the required view could be scored.

A real next step would validate required fields, evidence mode, and contiguous Score1…ScoreN ordering, repeat the borderline trial, and ask bicycle-maintenance experts and independent reviewers to test the wording.

Keep the evidence boundary: if the view or observation log does not support a score, mark the item unscorable: true, explain why, and route it for review. Do not turn missing evidence into a pass or a fail.

Prepared revision plan · not an executed transform or validation result

Step 8 · Ready

Readiness source: prepared v3 candidate plus human checklist; no Rubric Maker validation or domain-expert approval has occurred on this page.

Before approval, confirm

  • •
    Purpose

    Confirm the task and intended use are stated.

  • •
    Evidence

    Confirm each score points to something a reviewer can observe.

  • •
    Boundaries

    Confirm safety stops and unscorable-evidence handling are explicit.

  • •
    Version

    Confirm the candidate file and its history are identifiable.

  • •
    People

    Obtain sign-off from domain experts and intended reviewers.

Focus criterionv3 · candidate
Score1From review feedback
Stop for a damaged rim, frayed cable, tire contact, or a brake that does not engage.
Score2From review feedback
No stop condition is present, but at least one passing check is missing.
Score3Refined after T-04
Both pads contact the rim below the tire; after release, a visible gap returns at both pads and remains for one hand turn; the lever does not reach the handlebar.
After expert approval

Hand the rubric to the assessment workflow

The approved artifact can be uploaded or mapped for use in MAPLES and OASIS. That connected handoff is a separate operation and may invoke services; it is not performed by this guide.

See the MAPLES review tour

Ready does not mean frozen. Keep the rubric version tied to its evidence and revisit it when reviewers disagree, the task changes, or new failure cases appear.
1 of 8 Link to this step

Where this fits

Rubric Maker

The public Rubric Maker v0.1.0 plugin is an installable agent skill bundle for creating, importing, reviewing, transforming, validating, and testing OSCE rubric artifacts. The public package lists Codex CLI and Claude Code as supported runtimes and an academic-research-only license.

Wayfinder Rubric Studio

Wayfinder Rubric Studio is the MAPLES web experience for assisted rubric creation and refinement. Its visible compare-and-decide workflow is shown in the interns’ public article and YouTube video.

SimRubrics

SimRubrics is a separate research application for rubric development. It is part of the broader OASIS family, but this click-through is not a SimRubrics screen or a claim that MAPLES evolved from it.

Method and content notes

This click-through follows the public MAPLES Toolkit marketplace and Rubric Maker Skills catalog, with the workflow locked to the released rubric-maker-skill/v0.1.0 package. It mirrors the released skills’ sequence for faithful import, review, transform, dry-run grading, and dry-run evaluation; it does not run those skills or present unreleased capabilities.

All rubric wording, walkthrough suggestion labels, observations, scores, decisions, and revision notes on this page were newly authored for the bicycle example with Codex assistance. They contain no private rubric wording, private outputs, clinical data, learner data, or real-person data. They illustrate how to reason about a rubric; they are not outputs from a recorded Rubric Maker review, transform, or grading run, and they do not show that Rubric Maker validated this particular result.

Open the exact v0.1.0 release · Browse every OASIS tour · Return home

OASIS — Open Assessment and Scoring Infrastructure Stack

 
  • Technical report

  • Demo repository

  • arXiv preprint

  • Jamieson Lab at UT Southwestern

  • Andrew Jamieson faculty profile