Skip to content
Nitesh Tiwari

Independent portfolio project

AI Learner Diagnostic

An independent prototype for diagnosing a learner’s skill gaps and proposing the next step, designed around what happens when the AI is unsure or wrong.

Outcome
Independent prototype · deterministic demo, no real users or model results
Focus
AI productEvaluationHuman oversight
Evidence
Independent prototype

Independent portfolio project, not shipped at any employer. The interactive demo runs on deterministic logic so the experience is reliable and transparent. No real-user adoption, model accuracy or production results are claimed.

The 30-second version

Read the full story
  1. 1Problem

    Working out what a learner is missing and what they should do next is judgment-heavy work that teachers rarely have time to do for every student.

  2. 2Approach

    Split the job between rules, a model and the teacher; designed confidence, fallbacks and override into the UX; defined the evaluation and launch gate before any model work.

  3. 3Outcome

    A working, deterministic prototype of the product loop. No real-user adoption or model-performance results are claimed.

01

Why AI

Framing

Problem first, model second.

Working out what a learner is missing, and what they should do next, is judgment-heavy work. Teachers do it well and rarely have time to do it for every student. The gap is not content. It is diagnosis at scale.

It is also a problem where being wrong is costly in a quiet way. A learner sent down the wrong path doesn’t complain, they just stop. So the design question was never whether a model can do this. It was where a model should do it, and what happens when it is wrong.

The diagnostic loop

Product reasoning
  1. Assess

    Collect a small, representative set of learner evidence.

  2. Diagnose

    Identify the skill gaps the evidence actually supports.

  3. Explain

    Show the learner and educator why, with evidence before the verdict.

  4. Recommend

    Propose a learning path the educator can accept or override.

  5. Practice

    Generate targeted practice for the diagnosed gap.

  6. Evaluate

    Score the outcome against a rubric, not a demo prompt.

  7. Adapt

    Feed the result back into the next decision.

  8. ↺ back to step 1

A loop, not a one-shot answer. Each step is a product surface with its own failure mode, which is why explanation and evaluation are steps in the loop rather than afterthoughts.

02

Prototype

Interactive prototype

Run it. Then try to break its confidence.

Choose answers that ignore the learner signal and watch the output change. Confidence drops, the evidence trace shows why, and at low confidence the system stops guessing and hands the decision to the educator.

Learner Diagnostic

Demo mode · deterministic

Try it · 3 signals, ~1 minute

See the diagnostic loop, not just the architecture.

Answer three representative learner-signal questions. The prototype turns those signals into a transparent diagnosis, states how confident it is, and hands the final call to the educator.

This is a deterministic product prototype, not a claim of production model performance. The production architecture below would use an LLM, retrieval, evaluation data, and human controls.

03

Rules vs. model

The first decision

Where rules win, and where a model earns its place.

The first design decision was a split, not a model choice. Anything that has to be consistent, auditable or cheap stays deterministic. The model gets the parts that need judgment over messy evidence, and each of those parts has a defined fallback. The teacher keeps the final call.

Proposed production split

Product reasoning
  • Score answers

    Rules

    Deterministic and testable. There is nothing to guess.

  • Combine accuracy, hints and time into a readiness signal

    Rules

    The same input must give the same output, and a teacher must be able to check it.

  • Map a pattern of errors to a likely misconception

    Model

    Needs judgment across messy evidence. Rule-based in the prototype.

  • Explain the diagnosis to learner and teacher

    Model

    Language generation, grounded in the evidence trace.

  • Generate targeted practice

    Model + bank

    Variety, anchored to a curated, vetted practice bank.

  • Decide what happens at low confidence

    Rules

    A fallback has to be predictable.

  • Accept or change the plan

    Teacher

    Accountability stays with a person. Overrides are logged.

In the prototype every row runs on deterministic rules, so the demo is reliable and inspectable. This table is the proposed production split.

04

System design

System thinking

Where the model sits, and where it doesn’t.

A practical architecture combines structured learner signals, retrieval from a curated knowledge base, an LLM for the judgment-heavy steps, deterministic scoring wherever possible, and a feedback and evaluation loop.

Illustrative model
  1. 01 · Inputs

    • Structured learner signals
    • Answers, attempts, hints, time on task
  2. 02 · Grounding

    • Retrieval from a curated knowledge base
    • Skill map and practice bank
  3. 03 · Reasoning

    • LLM proposes diagnosis and next step
    • Deterministic scoring where possible
  4. 04 · Experience

    • Evidence, confidence and explanation
    • Educator accept or override
  5. 05 · Learning loop

    • Evaluation dataset and rubric
    • Overrides logged as feedback
The learning loop feeds overrides and evaluation results back into inputs. The model (highlighted) is one layer of five, and most of the product risk lives in the other four.
05

Failure modes

Failure modes

What happens when it’s wrong.

Every AI feature fails. The product decision is how: what the user sees, how the failure is detected, and what the system does instead. Designing these before the happy path keeps the demo honest.

Failure modes, designed before the happy path

Product reasoning
  1. 01

    Overconfident diagnosis

    Looks like: A firm verdict from two answers

    Caught by

    Too little evidence for the confidence claimed

    Product does instead

    State low confidence, hold the current path, ask for more evidence

  2. 02

    Hallucinated gap

    Looks like: A skill the learner was never tested on

    Caught by

    Every claim must trace to an answer in the evidence

    Product does instead

    Drop untraceable claims before anything is shown

  3. 03

    Contradictory signals

    Looks like: Fast, correct answers with heavy hint use

    Caught by

    Signals disagree beyond a set margin

    Product does instead

    Flag it for the teacher instead of picking one reading

  4. 04

    Discouraging language

    Looks like: “You are weak at fractions”

    Caught by

    Tone checks in the evaluation set

    Product does instead

    Describe the gap and the next step, never the learner

  5. 05

    Slow or failed model call

    Looks like: A learner waiting at a checkpoint

    Caught by

    Latency budget exceeded or timeout

    Product does instead

    Serve the rule-based next step and diagnose in the background

  6. 06

    Teacher disagrees

    Looks like: An override

    Caught by

    Every override is logged

    Product does instead

    The override wins, and becomes a new evaluation case

06

Evaluation

Evaluation

Evaluation is a product practice.

Before launch, evaluate diagnostic accuracy, recommendation relevance, groundedness, harmful or overconfident outputs, consistency, latency and cost, against a representative evaluation set rather than a few demo prompts.

Pre-launch evaluation rubric

Product reasoning
  • Diagnostic accuracy

    Does the diagnosis match what an expert educator would conclude from the same evidence?

    If it fails: Wrong path for the learner

  • Recommendation relevance

    Is the next step the most useful one for this gap, at this level?

    If it fails: Busywork and disengagement

  • Groundedness

    Is every claim traceable to learner evidence or curated content?

    If it fails: Hallucinated gaps

  • Overconfidence and harm

    Does it hedge when evidence is thin, and avoid discouraging language?

    If it fails: Loss of learner and educator trust

  • Consistency

    Do similar learners get similar diagnoses across runs?

    If it fails: Unpredictable experience

  • Latency and cost

    Is it fast and cheap enough to run at every checkpoint?

    If it fails: Unviable unit economics

07

Guardrails

Guardrails

Designing for uncertainty.

Show evidence where it exists, never present uncertainty as certainty, let a person override, and define what the system does when learner signals are thin or ambiguous.

  • Show the evidence

    Every diagnosis lists the signals it used, so it can be checked.

  • Honest confidence

    Uncertainty is stated, never styled away.

  • Human override

    The educator makes the final call; overrides are logged as feedback.

  • Defined failure state

    When signals are ambiguous, the system holds instead of guessing.

08

Launch criteria

Launch criteria

A gate, not a date.

Ship only when quality thresholds hold across representative cases and the AI experience measurably beats a credible non-AI baseline.

The first real test would be narrow: one subject, a few teachers, and the prototype’s diagnosis compared with the teacher’s own call on the same evidence. If it can’t agree with teachers often enough to save them time, it isn’t ready, however good the demo looks.

Launch gate: all four must hold

  1. Quality thresholds met across representative evaluation cases, not a handful of demo prompts
  2. No harmful or overconfident outputs in the failure-category review
  3. Latency and cost viable at every learner checkpoint
  4. Measurable improvement over a credible non-AI baseline

“Ship when it beats a credible non-AI baseline, not when the demo looks good.”

What this prototype is, and isn’t

It demonstrates the product loop and the decisions around it. The demo uses deterministic logic so it behaves the same way every time; a production version would connect the same flow to an LLM, curated retrieval, an evaluation set and human controls. It has no real users, and no adoption or model-performance results are claimed.