An AI assistant forchemistry teaching

A research project of the AG Chemiedidaktik at the Carl von Ossietzky Universität Oldenburg: Chemie.KI evaluates student answers based on the assessment rubric that you set and release – criterion by criterion, with partial credit instead of a blanket grade.

  • scientifically sound
  • for teachers
  • Carl von Ossietzky Universität Oldenburg

New: Chemie.KI has been extensively revised and expanded. The assessment now stands on a better foundation, you can now create your own tasks – and registration is temporarily open to all chemistry teachers.

The principle

How does Chemie.KI work?

For each chemistry task we store a fixed assessment rubric with individual criteria – for example: Was the technical term used correctly? Was the justification logically derived? A language model then checks the student answer criterion by criterion and awards partial credit instead of a blanket grade.

To make the AI judge like an experienced chemistry teacher, we additionally train it with thousands of real student answers rated by experts. This way the model also learns to handle incomplete or linguistically imperfect answers – and in initial evaluations achieves assessment quality close to that of human raters.

The feedback

Each point has an evidence passage

It is not the answer as a whole that is graded, but each criterion separately. Each awarded partial credit is accompanied by the passage in the response text on which it is based – color‑highlighted and traceable for learners and for you.

What a criterion does not cover remains visibly open: highlighted with dots, with a note about what would have been required for full credit. Only what the task actually requires is graded.


Try it yourself

Select a task, write an answer and have it assessed. The result will appear in the card next to it.

Explain why a sodium hydroxide solution reacts as basic. Also state how the pH value changes compared with pure water.

Acid–base · AFB II · 3 points

Explain why a sodium hydroxide solution reacts as basic. Also state how the pH value changes compared with pure water.

Sodium hydroxide dissociates in water into Na+ and OHK1. As a result the concentration of hydroxide ions increasesK2 above that of hydronium ions, therefore the pH value of the solution changesK3.

The colored highlight marks the passage on which the assessment of a criterion is based. Hover over a criterion below to emphasize its evidence passage.

  • K1

    The dissociation of the solid in water is correctly named with both ions.

    Dissociation described

    1/1
  • K2

    The excess of hydroxide ions compared with hydronium ions is given as the cause.

    Justified ion ratio

    1/1
  • K3

    It is mentioned that the pH value changes – but not in which direction compared with pure water.

    pH value given · Note: The pH value is above 7, the value of pure water

    0/1

Score achieved2 out of 3

Designing a task

From task to feedback

Four steps, from creating your own task to evaluating the entire class group.

  1. 1

    Set task and criteria

    You write the task and each assessment requirement on its own line with a point value. The total score follows automatically.

    teacher
  2. 2

    Release assessment rubric

    Chemie.KI generates a sample solution, partial credit rules and borderline cases from this. You read, correct and release them. Nothing is graded without release.

    teacher
  3. 3

    Edit test by code

    You bundle released tasks into a test. Your class group opens it with a five-digit code – no account, no name.

    Class group
  4. 4

    Feedback and analysis

    Each response receives partial credit, a justification per criterion and the supporting evidence passage. You see the class at a glance – including marked borderline cases.

    Both
View the four steps as a guided demo
As of today

Figures and facts

Four figures against which the project must be measured.

17.693 annotated cases

Student answers assessed by experts – the basis for every assessment.

κ = 0,87 Agreement with the gold standard

Human teachers achieve 0.90. See the study

without names Participation via code

Learners take tests via a five-digit code – without an account and without personal data.

Get started

Read all the way to the end? Great!

Then get started – it costs nothing. Create an account, release a task, let your class group participate using a code. The rest follows as you try it out.

free of charge as part of the research project no data from your class group required can be deleted at any time

Reliability

Control remains with you

A KI that grades performances must remain verifiable. Four provisions ensure this.

Nothing runs without release

A task is only graded after you have read and released the assessment rubric.

Chemie.KI suggests a model answer, partial credit rules and borderline cases – none of this becomes binding before you have confirmed it. Until then, the task remains invisible to your class group.

Your criteria remain your criteria

Number, order and point values are checked at every model call. Deviations lead to an abort – not to a silent correction.

The model may therefore assess, but not rewrite what is being assessed. A criterion you set at one point cannot suddenly be worth two.

Every assessment is recalculated

Partial credits may not exceed the maxima; the sum must be correct. The system checks both itself.

The calculation is arithmetic, not an assessment – so it does not belong in the language model, but before and after it.

Borderline cases come back to you

Uncertain cases are marked for review instead of being passed. Manually graded responses can be stored as a benchmark.

Over time, this changes not the assessment, but only the number of cases you still have to review yourself.

And what if the AI gets it wrong? You will see it: Each item of partial credit includes the evidence passage in the response text and a one-sentence explanation. An assessment you cannot understand is, in case of doubt, no assessment at all.

How well this works on average has been measured and published.

To the study
Tasks

Two ways to a graded task

You can use the reviewed pool or create your own tasks – both end up in the same test.

Ready prepared

Tasks from the pool

For 13 subject areas there are 65 reviewed tasks with ready assessment rubrics – supported by thousands of student responses graded by experts.

  • Select a task and add it directly to a test
  • View the assessment rubric and copy it for your own tasks
  • Ready to use without preparation
View demo
Self-created

Own tasks

Enter your own task description, set criteria, generate an assessment guide, review and release – for a topic from the existing collection or for your own topic.

  • You determine criteria and point values
  • Fine adjustments for operator, AFB and strictness possible
  • Own assessed answers can be stored as calibration
Create account

Both paths end the same way: You bundle released tasks into a test that your class group opens with a five-digit code – without an account and without names. Tasks taken from the existing collection can still be adapted beforehand; the assessment guide remains yours.

You can look around first and decide later – the demo does not require an account.

Guided demo
Figures and facts

What is behind the four figures

Where the figures come from, what they mean – and where their limits lie.

17.693 annotated cases

Real student responses, assessed by people

The figure is the current status of our database – not an estimate, but the collection counted when this page is accessed. Each case is an actually written response from chemistry lessons, assessed criterion by criterion by experts.

From this collection, the model learns what a response is worth in practice: with spelling mistakes, incomplete sentences, correct ideas in unfamiliar words. That is why it judges more leniently than a simple keyword search and more strictly than a charitable skim.

κ = 0,87 Agreement

Close to two teachers assessing the same thing

Cohens κ measures how closely two assessments agree – adjusted for the matches that pure chance would explain. 1.0 would be complete agreement, 0 would be guessing.

When two experienced teachers independently assess the same responses, they reach about 0.90 in our data collections. Chemie.KI reaches 0.87 against the gold standard. The gap is small, but it exists – and that is precisely why release remains with you and borderline cases are flagged.

a few seconds until feedback

Feedback while the question is still open

The assessment is generated during processing: submit, wait briefly, read the feedback. The exact duration depends on the length of the response and the workload – it takes seconds, not minutes.

The didactic difference lies not in the speed, but in the timing: Anyone who only finds out after a week what the problem was has long since forgotten their own reasoning. Anyone who finds out immediately can take it up again.

without names Participation via code

Learners do not need an account

A test is opened using a five-digit code. No account is created, no name is requested, and no email address is stored – the answer is recorded in the database without any person behind it.

For you, this means: you see the class group as a whole and each answer individually, but the association with a name is created only where you establish it yourself – in the classroom, not in the system.

Status and limitations: Current figures are counted afresh each time the page is loaded. The agreement comes from a data collection using a fixed set of tasks and does not automatically apply to every new task – for your own tasks, it is an indication, not a promise.

The method, data basis and analysis are described in detail.

The science behind it