Startup articles: launches, insights, stories

Token Cut - Startup logo and branding

Cut LLM Token Costs With Jev Model Routing

Founded year: 2026
Country: United Kingdom
Funding rounds: Not set
Total funding amount: $1000000

Description

Most AI products have a surprisingly expensive default:

Send everything to the same powerful model.

It works, but it’s wasteful.

A simple classification request doesn’t need the same model as a difficult coding problem. Summarizing an email doesn’t require the same level of intelligence as solving a multi-step reasoning task.

That’s the idea behind Token Cut.

Token Cut evaluates each request and routes it to the lowest-cost model capable of handling the task well.

How it works

Every request first goes through Jev, a small decision model that identifies the task and estimates the level of capability it requires.

Token Cut then evaluates available routes based on task type, model capability, modality, context requirements, budget, routing confidence, and provider pricing.

Simple request → cheaper model.

Reasoning or coding task → stronger model.

Difficult task → frontier model.

The goal isn’t to always use the cheapest model. It’s to stop using expensive models when you don’t need them.

One API instead of hardcoding models

Without routing, you might end up writing logic like:

if task == "simple": use_model_a()

Then a new model launches. Pricing changes. Another provider becomes cheaper. Soon you’re maintaining a growing list of routing rules.

With Token Cut, your application sends the request to one API and Token Cut handles the model selection.

We’re integrating providers including OpenRouter, Replicate, fal.ai, and Kie.ai, giving the routing layer access to a much larger pool of models and pricing options.

Optimize cost before spending it

A lot of LLM cost tools focus on observability.

They tell you which models consumed tokens and how much money you spent.

That’s useful, but the money is already gone.

Token Cut approaches the problem earlier.

Before sending the request, it asks:

Does this task actually need an expensive model?

If a cheaper model can handle it, Token Cut routes the request there. If the task requires more capability, it can use a stronger model.

Why we’re building it

There are more AI models and providers than ever.

That’s great for developers, but choosing between them is becoming an engineering problem of its own.

Capabilities change. Prices change. Context limits change. New models appear every week.

We don’t think developers should have to constantly think about individual model names.

You should be able to define what matters — quality, capability, latency, and budget — and let the routing layer figure out which model makes sense for each request.

That’s what we’re building with Token Cut.

Try it: https://www.tokencut.lol

It’s still early, and I’d love feedback from developers already spending real money on LLM APIs.

How are you handling model selection today?

Have a startup you want to promote?

Feature your startup and reach thousands of entrepreneurs and investors

Feature My Startup

Related startups: