I’ve been working with the new Engineering Notebook Rubric (v1.2) and wanted to share some concerns about ambiguous language that seems to be causing confusion for students and inconsistency for judges after being the judge advisor for our first Level Up event.
My goal here is not to criticise the intent of the rubric, but to highlight areas where clearer wording or more specific exemplars could improve both student understanding and judging efficiency.
1. “Complete and professionally presented” / “Highly professional”
Under Cover and Team Information, “complete and professionally presented” is not defined. Different judges can interpret “complete” very differently (e.g., does this include photos, contact details, school logos, etc.), which may lead to inconsistent scoring.
Similarly, Professional Presentation – “Highly professional and polished” is challenging when we’re talking about students, many of whom are in middle school. While some subjectivity is necessary in judging, leaving this level of professionalism undefined slows down the process and can create large differences between judging teams. More concrete descriptors (e.g., “legible handwriting, consistent formatting, minimal corrections, appropriate language”) would help.
2. Problem Definition
The Problem Definition criterion is currently phrased as “Thorough analysis and constraints,” but it doesn’t specify what “problem” is being defined. Is this the game challenge, the robot design problem, or a specific subsystem issue?
In practice, I see two distinct areas here:
-
Game analysis and constraints based on the manual.
-
Robot design problems and constraints (e.g., mechanisms, strategy, field navigation).
I would recommend splitting this into separate criteria or explicitly stating that teams should define both the game problem and the robot design problem, so students and judges have a clear target.
3. Design and Coding Decisions vs Brainstorming
Brainstorming and Design and Coding Decisions feel like overlapping criteria. Brainstorming already implies generating options; design and coding decisions cover selecting among those options.
A combined criterion such as “Brainstorming and Decision Making” could make the expectation clearer: teams should show a range of ideas and then explain how they evaluated those ideas and chose the best solution, supported by evidence.
4. CAD/Build/Code Documentation
The CAD/Build/Code Documentation row asks for “Comprehensive records of builds and/or coding projects,” but students (especially IQ teams) are unsure what “comprehensive” looks like here.
For example:
-
Does “comprehensive” require CAD for all teams, or are hand sketches acceptable at IQ level?
-
For coding, is the expectation pseudocode, flowcharts, screenshots, or printed code snippets with commentary?
More explicit examples or tiered expectations (IQ vs V5) would help students document appropriately and reduce guesswork.
5. Calculations
The Calculations criterion (“Extensive quantitative analysis / Appropriate calculations”) is one of the most ambiguous areas. Judges are left to decide what counts—gear ratios, torque, velocity, scoring projections, statistical analysis of test data, etc.
Because this is so open-ended, I’m already seeing inconsistent marking practices across events. It would be helpful to list typical calculation types or link to guidance so teams understand what’s expected at different levels.
6. Data Collection and Analysis
In the Testing, Analysis, and Iteration section, Data Collection and Analysis & Improvements are separated. Conceptually, this makes sense, but in the context of teaching the iterative design process, these are tightly coupled.
I’d suggest either:
-
Keeping them separate but explicitly describing the relationship (collect data → analyse → iterate), or
-
Combining them into a single criterion such as “Data Collection and Analysis,” emphasising how teams gather data and use it to inform iteration.
7. Team Roles – “Roles and responsibilities evolved”
Under Team Roles, the exemplary descriptor “Roles and responsibilities evolved” is difficult for students to interpret. “Evolved” is vague: does it mean roles changed over time, became more specialised, or were refined based on team needs?
Clearer language like “Roles and responsibilities are clearly defined, documented, and refined over the season based on team experience” would help teams know what evidence to include.
8. Competition Reflections and Failure Analysis
The Competition Reflections & Failure Analysis criterion is excellent for encouraging reflection, but it raises equity issues between teams. A student at their third event has significantly more competition experience to analyse than a team at their first event.
As written, teams without prior competition experience are naturally disadvantaged because they lack that specific type of failure analysis. Broadening the criterion to include failure analysis related to robot design, testing, and practice sessions (not only competitions) would allow all teams to demonstrate reflective thinking and improvement.
9. Continuity vs Pagination and Chronology
Finally, Continuity under “Completeness and Season Narrative” overlaps heavily with Pagination and Chronology in the first section.
Both criteria essentially ask whether the notebook is:
-
Easy to follow.
-
Logically ordered.
-
Telling a coherent story over time.
It would be helpful to distinguish these more clearly (e.g., one focused purely on physical organisation/pagination, the other on narrative coherence) or consolidate them to reduce redundancy.
I’m sharing these thoughts in the hope that we can refine the rubric language so it:
-
Gives students clearer guidance on what to document.
-
Supports more consistent judging across events.
-
Keeps the judging process efficient while still honouring the complexity of engineering design.
-
I’d be very interested to hear how other EPs, judges, and educators are interpreting these criteria and whether you’re seeing similar issues at your events.