AI strategy & applied intelligence

The future belongs to differentiated intelligence.

Foundation models are becoming a commodity. What creates enduring value is not the base model. It is how your organization teaches it to think, decide, and act in the context of your own business.

I help startups and enterprises build that intelligence, and keep it inside their own trust boundary. Advisory, hands-on build, and a diagnostic that tells you where your judgment actually lives.

Anub Sinha · engineering leader, builder at Opscale and LegalStreet, author of the Applied series.

The argument

Three beats

The strategy question for the next few years is not which frontier model you use. It is what the one model only you could have built is, and whether you get to keep it.

01

Rented intelligence is not an edge

If my competitors and I are renting the same intelligence, then none of us has a durable edge. We are all standing on the same floor, paying the same rent, arriving at the same answers. The frontier model is table stakes. It was never the moat.

02

You cannot prompt taste

Real judgment comes from experience, and experts struggle to articulate it even to other humans. A prompt can describe judgment. It cannot transfer it. To get judgment into a model you have to train it in, on examples only your people could have labelled.

03

Your corrections are the asset

Every trace, trajectory, and correction your experts produce is the raw material of differentiated intelligence. If it flows out to the provider, you are paying to build an asset your competitors will rent back from them later. That is the leak worth closing.

Renting

Your experts do the hard work. The learning leaves with the provider, and reaches everyone who pays the same subscription.

YOUR ORGANIZATION Your experts Traces, corrections, trajectories Shared frontier model Rival A Rival B Your edge, redistributed.

Owning

The same work, kept. Every correction compounds into a model that gets sharper at your business, and nobody else's.

YOUR TRUST BOUNDARY Your experts Traces, corrections, trajectories Your tuned model Nothing crosses the line. The loop compounds.

This is not a forecast

The evidence

Bridgewater and Thinking Machines took a task investment professionals do every day, and tested whether a tuned open model could beat the frontier at it.

84.7% Accuracy from a fine-tuned open model, against 78.2% from the best frontier model with expert prompting.
~30% Fewer mistakes than the frontier model. The difference between below the trust bar and above it.
~14× Cheaper to run at inference, because the specialist model is smaller than the generalist.

Out of the box, frontier models scored around 50% on the task. Expert prompt engineering pushed them to roughly 78%, still under the 80% bar a professional needs before they will trust the output. The win did not come from prompting harder. It came from training judgment in, using expert-labelled data the firm already had.

That is the whole thesis in one experiment. Smaller model, your data, your taste, beating the giant that serves everybody.

Source: Thinking Machines with Bridgewater, Learning to Replicate Expert Judgment in Financial Tasks (2026).

“Outperforming the market is hard. When every investor has access to the same sources of public information, alpha must come from unique insight built on taste and judgment.” Thinking Machines × Bridgewater. The same logic holds for every industry that has just been handed the same API.

How I work with you

Four engagements

Most engagements begin with the diagnostic, because the expensive mistakes happen before anybody trains anything.

Intelligence Diagnostic

Start here

Two to three weeks, working with your leadership and the people who actually make the calls. We map where proprietary judgment gets created in your business, what is currently leaving your trust boundary, which of it is genuinely defensible, and what is realistically tunable today. You finish with a clear picture and a ranked plan, whether or not you keep working with me.

  • A map of where judgment is made, and by whom
  • An audit of what leaves your boundary today, including vendor terms
  • A ranked shortlist of candidate tasks worth tuning for
  • Honest calls on what you should keep renting

You get: a written diagnostic, a ranked roadmap, and a decision on build versus rent that survives scrutiny from your board.

Advisory

Ongoing

A standing relationship with founders, CTOs, and heads of AI. The questions that do not fit into a project: build versus rent, architecture, what to keep in house, how to structure the data and the team, what to insist on in a vendor contract.

  • Regular sessions with the leadership team
  • Review of architecture and vendor terms
  • On call for the decisions that are hard to reverse

You get: a senior second opinion before the costly commitments, not after.

Hands-on build

Project

The work itself, with your engineers rather than around them. Instrumenting the traces, standing up the expert labelling loop, fine-tuning and evaluating against your experts' bar, and deploying it where your data already lives.

  • Data capture and expert labelling pipelines
  • Fine-tuning, distillation, and evaluation harnesses
  • Deployment inside your trust boundary
  • Handover your team can actually run

You get: a working model, the loop that keeps improving it, and a team who can operate both.

Workshops & talks

Half or full day

For leadership teams and boards who need to reach the same understanding at the same time. The argument, the evidence, and a working session applying it to your own business rather than a generic case study.

  • Executive briefing on differentiated intelligence
  • Working session on your own candidate tasks
  • Conference and offsite keynotes

You get: a leadership team that stops arguing past each other about AI strategy.

The compounding loop

The method

Differentiated intelligence is not a project you finish. It is a loop you close, and then keep turning.


01

Locate

Find where proprietary judgment actually gets made. It is rarely where the org chart says it is, and it is never in the parts already easy to automate.

02

Capture

Instrument it. Traces, corrections, and the reasoning behind the calls become a dataset instead of evaporating into chat logs you do not own.

03

Train

Tune a model on it and hold it to your experts' bar, not a public benchmark. The bar is whatever score makes your people trust the output.

04

Compound

Feed every correction back inside the boundary. The gap between you and anyone renting the same base model widens each quarter you keep turning it.

The loop only compounds if it closes inside your boundary. If step four leaks, you are running a very expensive training program for the whole market.

Whether this is for you

Plainly

I take a small number of engagements at a time, so it is worth being direct about where this works and where it does not.

A good fit

  • You have proprietary data, or experts making judgment calls that are hard to write down
  • There is a real decision in your business that AI keeps getting almost right
  • You are at enough scale that a few points of accuracy change the economics
  • You are uneasy about what your vendor terms let providers learn from you
  • You want your team to own the capability afterwards, not depend on me

Not a good fit

  • You want AI added to the product without a specific decision it should improve
  • There is no proprietary data and no expert bench to learn from
  • You are looking for a thin wrapper over an API, shipped this month
  • The goal is a demo for a fundraise rather than a capability that lasts
  • Nobody senior is willing to spend time on the labelling and the standard

Who you would be working with

I am Anub Sinha. I build software and write about the ideas underneath it. I am currently building at Opscale and LegalStreet, and I write the Applied series, two online books that take one idea at a time and chase it from first principle to the machine it ends up inside.

That is the same instinct I bring to this work. Most AI strategy advice stops at the level of the slide. The useful part is one layer down: what the data actually looks like, where judgment actually gets made, what a training run actually costs, and which of it your team can actually run once I leave.

I have been making the case for differentiated intelligence since well before it had evidence behind it. Now it has evidence.

What is the one model only you could build?

If you can answer that, you probably do not need me. If the question makes you uncomfortable, that is the conversation worth having.

Engagements are senior, scoped, and priced for the decisions they change. I run a small number at a time.