Skip to content
Nitesh Tiwari
Back to home

AI Lab

AI product judgment: problem first, model second.

Strong AI product judgment; production evidence in progress. No evaluation has been run.

  • Implemented

    Working code in the repository

  • Designed

    A written design, not yet code

  • Planned

    Scoped, not designed in detail

  • Needs input

    Waits on a decision, data or access

Professional context: at Edfora the adaptive engine used 3PL Item Response Theory, a statistical model rather than an LLM. It is a personalization case, not an AI-model claim.

In the field

When the AI option didn't ship

Edfora · myPAT and Glorifire doubt resolution

  1. Signal

    An AI auto-resolver was fast and highly scalable; the risk was incorrect answers damaging academic trust.
  2. Decision

    A hybrid of guided hints, verified peer solutions and SME escalation, behind a 90% accuracy circuit-breaker with a rollback.
  3. Outcome

    From a ~24 h reported peak to a <15 min median for repetitive doubts in the pilot; >90% accuracy held on the hybrid. AI was evaluated but not shipped as the first solution.

Read the case

Current build

Independent prototype · deterministic baseline

AI Learner Diagnostic

A deterministic, rules-only prototype; not an LLM. The baseline, the override and the evaluation harness are built. No evaluation has been run, and there are no real users or model results.

  • ProblemDesigned
  • Why AIDesigned
  • BaselineImplemented
  • Human overrideImplemented
  • Output schemaImplemented
  • Evaluation harnessImplemented
  • Evaluation casesDesigned
  • ArchitectureDesigned
  • Failure taxonomyDesigned
  • Launch gateDesigned
  • Model decisionPlanned
  • Context and retrievalPlanned
  • MonitoringPlanned
  • Model-based diagnosisPlanned
  • Educator label reviewNeeds input
  • Latency and cost budgetsNeeds input
  • API accessNeeds input

Open the build

What I hold every AI build to

  1. 01

    Rules first, a model where it earns it

    Much of a learning product can run on deterministic logic. A model belongs where judgment is needed and a wrong answer is recoverable.

  2. 02

    Evaluation before scale

    A representative test set, a rubric, named failure categories and a launch threshold, agreed before anyone argues about the demo.

  3. 03

    Uncertainty is a UX problem

    Show the evidence, say how confident the system is, define what happens when it isn't, and let a person overrule it.

  4. 04

    Cost and latency are product constraints

    If it can't run at every checkpoint at an acceptable speed and cost, it is a demo feature, not a product.