Insights

AI Made Output Cheap. Judgment Is the New Founder Signal — KellyOnTech

The Judgment Trail: Output Is Cheap Now. Judgment Isn’t

Canva now expects backend, frontend, and machine-learning candidates to use AI during technical interviews. The real test is whether they can stay in control of the AI while they use it.

Candidates with less AI experience struggled not because they lacked coding ability, but because they lacked the judgment to guide AI effectively and catch its mistakes. Two candidates can produce identical output. One questions underlying assumptions, tests the code, and catches subtle errors. The other accepts what the model hands back.

Same output. Very different judgment.

The old model of evaluation is breaking down

Meta has been piloting AI-enabled coding interviews since 2025 on the logic that they reflect how engineers actually work—and make AI-based cheating beside the point, since the AI is already in the room. CoderPad’s customers alone have run more than 35,000 AI-assisted interviews.

And recent Harvard Business Review research analyzing thousands of screening sessions reached a similar conclusion:

When competence is this easy to simulate, judgment becomes the scarce signal.

The old model of evaluation is breaking down Mans International
The old model of evaluation is breaking down

This shift isn’t limited to technical hiring—it applies directly to how founders make strategic decisions and how venture investors evaluate startups. When AI makes execution outputs instant and cheap, a founder’s pitch deck or product prototype no longer proves strategic insight. Real differentiation lies in how accurately a leader isolates untested market risk and stress-tests core assumptions before burning capital.

AI is making output easier and cheaper to produce. What’s left to differentiate is the judgment behind it. Traditional evaluation looks at endpoints:

Credentials → Output → Outcome

But when output is cheap, the artifact stops telling us much about the capability behind it. We have to look upstream:

Context → Assumptions → Reasoning → Decision → Outcome → Learning

From Output to Judgment Mans International
From Output to Judgment

I learned this a decade before generative AI made it obvious.

A ten-year lesson, finally measurable

Early in my ten years as a founder, I worked with a Canadian data-services company to explore market opportunities and secure funding.

The technology was strong—its models could clean bank card data, detect fraud, and significantly reduce acquisition and retention costs. The company had senior banking relationships, a clear commercial case, and a secured pilot.

The pilot still failed.

A ten-year lesson, finally measurable Mans International
A ten-year lesson, finally measurable

Not because the technology didn’t work, but because bank data was fragmented across departments, privacy restrictions complicated access, and internal teams resisted an outside solution that threatened legacy systems. Crucially, the startup lacked the capital to survive a multi-year enterprise sales cycle.

We had confused functional capability with systemic viability. We proved the technology could clean data, but we ignored the complex operational system around it.

If we had maintained a Judgment Trail at the time, our focus wouldn’t have been on tracking code completion or pilot deliverables (output); it would have been on tracking our untested assumptions around enterprise adoption speed and internal stakeholder resistance (judgment).

What That Judgment Trail Might Have Looked Like:

  • Core Untested Assumption: Executive sponsorship and a signed pilot guarantee fast, cross-departmental data access before our runway runs out.
  • Operational Risk to Test: Required data is trapped in fragmented legacy systems, stalling the pilot behind multiple, disjointed legal, security, and IT approvals.
  • Validation Trigger / Pivot Rule: If we cannot secure a named data owner, a clear approval path, and a committed budget within the initial pilot window, then we immediately pause custom development and pivot to customer segments with simpler integration cycles.

Building your own Judgment Trail

Founders keep diaries. Founders keep to-do lists. Very few keep a record of why—the core assumptions they walked in with, what evidence changed their mind, and what they would execute differently now.

A diary tells you how a decision felt. A Judgment Trail tells you how it was constructed. Because a founder’s decisions, tracked over time, are the clearest evidence of how they lead, the Trail isn’t really a record of the business—it’s a record of leadership quality.

Building your own Judgment Trail Mans International
Building your own Judgment Trail

Once judgment becomes valuable, people learn to perform it. Two specific diagnostic questions cut through surface-level narratives faster than anything else.

  1. Test the trade-off behind the decision: “You had two viable options and chose A. Why was A better than B?”

    Good judgment makes the trade-off explicit. Weak judgment often retreats to vague answers such as “it felt right.”
  2. Test what they challenged before being asked: “What assumption or edge case did you personally test before anyone prompted you?”

    What someone chooses to verify on their own often reveals more about their judgment than a polished explanation after the fact.

Applying this framework in practice requires adapting it to your role:

For founders: Skip the heavy decision register—it won’t survive a busy execution week. Focus on lightweight logging: capture the initial context, 2–3 core unproven assumptions, and the specific triggers that would force a strategy pivot.

For investors: Look beyond the polished narrative deck. Ask for the team’s Judgment Trail to evaluate how they process dynamic market feedback and stress-test their own assumptions over time.

Shift your team from evaluating output to evaluating judgment. Contact us and I’ll send you the Mans International Judgment Trail template I use with portfolio founders.

Share: