Skip to content

AI-era talent
evaluated in a new way

Don't measure tool fluency — measure the thinking: how someone explains a problem to AI, validates the output, and chooses the right tool. Beyond the deliverable, how they solved it is captured automatically, so reviewers can see the whole process at a glance.

Probe candidate workspace — task start, web IDE solving, submit preview, completion

Talent in the AI era delivers different results under the same conditions

Handing everyone the same AI doesn't make the output the same. What sets them apart isn't the tool — it's the competency to operate it.

So, what are you looking at to hire AI talent today?

What makes a good AI hire, what 'using AI well' actually means, how much AI literacy the people evaluating themselves should have — none of it is clear.

Recruiter A

No AI hiring profile

Because each company and role needs different AI abilities, there's no agreed-upon, universal definition of an AI hire.

Recruiter B

No yardstick for 'good with AI'

Shipping code fast isn't the same as using AI well. You need a measure of what to look for, and how far it should go.

Interviewer C

I'm a reviewer, but I barely know AI

If the people evaluating aren't comfortable with AI tools and environments, they can't read the real gap between candidates.

Defining AI talent, setting the bar, automating assessment — all in one module

Probe gives recruiters an AI hiring standard, co-authors the problems with them, and runs an automated evaluation.

Usecase

KRAFTONCofa

The future of hiring evaluation,
built with an AI-first company

Probe is in production as Krafton's standard hiring assessment. From researchers and engineers to non-developer and back-office roles, Probe delivers trusted tests and assessment reports aligned to Krafton's Vision & Value.

AI-native hiring know-how

Krafton's criteria and methods for hiring at the AI frontier — packaged in.

In production, expanding

Krafton itself is rolling AI capability assessment across more and more roles.

Our Standard

We define AI-native competency, and build a new standard of evaluation

The AI-native competency companies want is the ability to produce great work by operating AI's strengths and limits within constrained resources. We evaluate it along three axes: commanding the artifact (technique), steering the process (intent), and understanding the result (cognition).

What is AI-native competency?

  1. 01

    Technical control (Technique)

    Is what the AI produced well put together under your own command — running cleanly, maintainably structured, and free of leaked secrets?

  2. 02

    Intent management

    Do you steer the AI — defining goals and constraints, decomposing and delegating work, and controlling results through verification loops?

  3. 03

    Cognition management

    Can you explain the AI's work grounded in your own artifact — why this structure, how it runs, and where it breaks?

Evaluation areas

We evaluate on technique, intent, and cognition

Each competency is scored against a detailed rubric. Every criterion gets a score, with met and unmet examples presented alongside the prompts and source code behind them — giving evaluators a signal they can trust.

Technical controlscored on the artifact

Functionality & execution

Does it run cleanly in a real environment?

Isolation & structure

Is it built on a maintainable, extensible structure?

Security hygiene

Do secrets and sensitive data stay out of code and logs?

Applied to every answer: no contradictions with the actual submission · grounded in their own artifact, not generalities · nothing fabricated · reasoning from evidence, not authority.

How it works

Candidates reason through the problem, and companies evaluate that reasoning

Candidates solve the task in a real environment, then answer questions about what they built. Both behavior (intent) and answers (cognition) are evaluated from real records.

  1. 1

    Start in a real environment

    Launch instantly in a browser IDE. Real APIs and real compute — not mocks — within a time and token budget.

  2. 2

    Artifact & process, auto-collected

    Source code, along with AI collaboration records like the transcript, prompts, tool configuration, and verification runs, is collected automatically. Candidates never write a separate report.

  3. 3

    Post-task Q&A

    Post-task questions span the artifact, intent, and cognition, and candidates must be able to explain their own work.

  4. 4

    Automated scoring · interview report

    Auto-scored against per-competency rubrics, producing a discriminating per-candidate report for the interview panel. Every criterion defines its grey area and met/unmet examples, so scoring stays reproducible.

Signals that count as met

  • Prompts with explicit done conditions & acceptance criteria
  • Fail → fix → re-pass through the same verification
  • Recorded rationale and rejected alternatives
  • Intent docs a new session can inherit (CLAUDE.md)
  • Answers grounded in their own artifact — matching what was submitted

Signals that count as unmet

  • Wholesale delegation like "just make it work"
  • Accepting AI's 'done' report without verification
  • Answers filled with textbook generalities
  • Explanations citing files or APIs that don't exist
  • Signal drowned by tool & MCP sprawl

Same starting line for everyone
and we focus on the problem-solving process

We place no caps on model power or scope — candidates get the same freedom they'd have in real work.