Evidence-based interview assessment for AI/ML engineers

How it works

An AI interview evaluator that refuses to judge what it has not tested

Every claim it makes is traceable to something you said. Everything it did not test is labelled untested, not weak.

Start free diagnostic

15 free minutes every week

  • Evidence-backed
  • No scores
  • Disputable claims

Coverage

3 of 6 areas checked

  • Prompt groundingChecked in this session
  • Re-rankingChecked in this session
  • Failure triageChecked in this session
  • Chunking strategyNot yet tested
  • Embedding choiceNot yet tested
  • Eval harness designNot yet tested

The problem with AI evaluation is confident nonsense.

It is trivial to have a language model listen to an answer and emit "78% — strong on fundamentals, weak on evaluation." It is also close to meaningless: the number has no basis you can inspect, the confidence is unearned, and the areas the conversation never reached get quietly folded into the same figure. An evaluator that will not tell you what it did not test is not evaluating, it is guessing with a decimal point.

Fifteen minutes later, you have three things.

  • A map of what was actually checked
  • A gap traced back to what you said
  • One thing to study tonight
See what each one looks like

Three rules the evaluation follows

These are product constraints, not preferences. They are why the output looks sparser than a scored report and why it is worth more.

  • Every gap claim carries the reasoning that produced it, drawn from the session
  • Concepts the conversation never reached are labelled "not yet tested", never as weaknesses
  • One recommendation at a time — a list of twelve is a way of avoiding a judgement

Claims you can reject

Each gap on the result page can be disputed. That is a real affordance, not a gesture: an evaluation you cannot argue with is one you have to take on faith, and taking model output on faith is exactly the failure mode this is built against.

Categorical judgements, not fabricated precision

Answers are assessed on correctness, depth, reasoning, and communication using categorical levels rather than continuous scores. A percentage implies a measurement precision that a fifteen-minute conversation cannot support, so the product does not render one anywhere.

Start free diagnostic

Nothing to install. Talk it through, keep the map.

Before you start

Why does it not give a score?
Because a score from a short adaptive conversation implies precision the evidence cannot support, and it hides the difference between "tested and shaky" and "never tested at all". The output separates those two explicitly instead.
How do I know the evaluation is not made up?
Every gap is shown with the reasoning that produced it — what you said, and what that suggested had not been tested. If the basis does not hold, you can dispute the claim on the result page.
What does it evaluate?
Explanations you give inside AI/ML engineering: ML, LLM, Applied AI, AI infrastructure, and MLOps. Coding throughput, delivery, and behavioural signal are not evaluated, and the evaluator reports them as out of scope instead of inferring them.
What does using the evaluator cost?
15 minutes of evaluation are free weekly and renew on that cycle, with no card requested. When one evaluation a week is not enough, Individual access is $29 a month.

Find the concept your prep has not tested.

15 free minutes every week. No card. One gap map, one next action.

Start free diagnostic

Related