JEV - Past that line, I can no longer tell who is smarter

The line is placed between Opus 4.5 and 4.6: most people can no longer pose a task beyond the model. Past it, what can still be told apart is the product definition. Jev emits a probability, not the next stretch of smarter chat.

Co-author: folkbench.com (Folkbench is a reviewable evaluation and selection platform for AI APIs, model services, and related sites. It draws on published service information, prices, availability, latency, and evidence windows to help users compare and choose a path.)

Ten pieces in this series:

The contest among foundation models is in its third year. Most of those still at the table have crossed a line. The line is not low.

My definition of it is one sentence. Ninety percent of users can no longer pose a task outside a model's reach. The work is still there. People do not stop asking. Those asks, the models can take. Replying to mail, collecting minutes, fixing one error, turning a pile of material into a few items: I stopped using these to rank them a long time ago. They sit inside the line. I place the landing between Opus 4.5 and 4.6. Above that, it blurs into one patch.

The two years before, I still chased releases and switched in to compare, seriously, who was stronger. This year I stopped comparing. The boards are still updating. The work in my hands has not gained another tier.

After the switch, only a feeling is left

Before the line, I could tell. Whether a long piece would come apart, I knew by switching models. Whether earlier conditions were still there once the steps piled up, I could feel within the week. After the line, I cannot feel the difference in capability. Switching models often leaves only an impression that it got a little better. Better at what, I cannot say.

The same kind of work needed my rework last month, and this month it goes through. Meeting notes go in, and the things to do come out. Two vendors come out a line or two apart, and both look usable to me. Code is the same. Paste an error, and both can give a change. Sometimes I feel one of them is cleaner, with a smaller diff. I cannot point to the update where they pulled apart. If I have to pick, I pick the smaller change. There is no day in the middle I can point at and say, from here it could do this. The increments have kept coming. They are too fine for my days to remember.

So I switched models less for this. Who is smarter at the next tier, I cannot tell.

These I can still tell apart

The first thing I can feel is the shape of what it emits. What emits text, I treat as someone to talk with. I hand it a messy record and have it collect the few things to do. I read it once, change a few words, and throw it back if I am not satisfied. That round trip is familiar. Voice is another path. I speak to it, and it speaks back. I do not use it much. I turn it on while I am on the way. The moment I need to change words, I go back to text I can copy. I recognize this shape.

What emits actions is more different. It operates a computer and clicks a flow through on its own. I watch its hands. Click the wrong place, and however complete the words after that are, they do not help. There is another kind that only gives a probability. Jev is that kind. It does not emit text, and it does not operate a computer. I have met all of these. I can tell them apart because the way I take over is different.

Temper and taste also separate people. Hand over the same material, and some write it very full, adding a sentence at the edge, afraid I will not understand. Others stop here and leave the judgment to a person. Who I keep often depends on whether that temper fits my hand. The one that stops in the right place, I keep in the window I open every day. A higher score beside it, and I still rarely switch windows for the score.

Cost and latency are felt together. Which stretch of the cost curve it sits on, and which stretch of the latency curve, two uses are enough to know. Chatting with a person, I can stand waiting a little longer. A person has to pause anyway. Put it on a flow that runs again and again, and waiting becomes a blockage. Slow enough that I can go pour a glass of water, and it only belongs in a conversation. Expensive enough that I do not dare call it once more, and it cannot enter the steps that are small and dense. If it is cheap, and it comes back fast, then I dare to spread it out. If what comes back is a passage I have to finish reading, cheap does not help. I still cannot put it in a loop.

Refusal is also a difference. What it refuses, and how it refuses, shows after a while of use. Some refuse directly, in very few words. I know the path is closed, I change the ask, and one changed sentence is enough to continue. Some do not say so. They walk around the thing and hand back work that has gone crooked. I will still think I did not say it clearly, and grind another round on a bad result. What I look at now is whether I dare to hand this work to it next time.

A new model states its place first

GPT-6 made the world recognize computer use again. Operating a computer was put back at the front of a release. What I care about is not that models could use a mouse before. Past the line, that is no longer new. I used to watch a release for another tier of height. This time I changed the question. A company releasing a new model has to think about its own niche first: who it is for, and what it is particularly good at. It does not have to stand up and say that it is the next, smarter tier.

Jev is an extreme I have actually touched recently. It does not compete with Opus at writing. Writing, it does not take. What it occupies is the joint. The material reaches the point where a decision has to be made, the question is closed, and the options were given in advance. What comes back is a probability. It comes back at once. On the chat side I still switch to another window and wait. Here I do not have to switch away. The result is already back. This kind of output is not charged. What I throw over is not "write me a paragraph". It is a judgment already circled: does this path go, and how sure is it. After the probability comes back, what takes over is a program, not me sitting there to finish reading and then decide. I have no desire to chat with it. The place is not reserved for chat. Chat models can go further up, and what they emit is still text. Jev removes the step of generating text. This is not smarter chat. It is another kind of product.

Between model companies, I now look at the route, and at whether the product definition can be carried through. How a model is designed now comes down to those two. It is not the first name on a board that decides. Names on the board change after a while. What a model intends to emit, once the shape is fixed, the training and the interface follow, and the pricing follows too.

Other routes, this piece only carries one sentence. Some push the engineering to an extreme. Some turn and build a vertical suite. The rest simply do not speak, and keep pushing a general model. The particular suites from DeepSeek and Kimi are not unfolded here. They are written in JEV - A stripped-down model, a suite, and one that does not speak.

Past that line, I can no longer tell who is smarter. When I switch, it can still feel a little better. That feeling is not enough for me to choose with. What I can tell apart is the product definition. Do you emit text, emit actions, or emit only a probability.

Sources

Series