Bloom

가장 활발하고 솔직한 AI 커뮤니티

000
KoEn

Price per token is the wrong unit now

What was discussed

Price per token is the wrong unit now

Two hundred seventy-one people signed up and the fifth floor of D.Camp Mapo filled up. The demographics from the registration form explained the topic on their own. Eighty-eight percent said they already look at benchmarks when choosing a model, and 36 percent train models or ship them in products themselves. Two hundred twenty-seven wrote in questions, and how scores are built topped the list at 55.

Artificial Analysis evaluates the whole stack, from models through inference to the hardware layer. What George said he is proudest of is that companies cite their charts not to show a single intelligence-index bar but to show the tradeoff a developer has to accept. That is why the charts are all 2x2. OpenAI used intelligence and cost; Jensen Huang used intelligence and speed.

His opening framing landed. We mostly use models as agents now, and it is worth pinning down where the magic comes from. A harness provides tools like code execution and web search. What drives what is possible is still the model. As evidence he showed the coding agent index. A weak model on a good harness did not beat a good model on an ordinary one.

The intelligence index is led by the newest models from Anthropic and OpenAI. What he wanted to talk about, he said, is why reasons remain to use other models. For a company, using a top model for every task could bankrupt it. The section on Korea drew the most nods. Over the past year Korea has risen to the top three, behind the US and China. Countries like Canada or Spain have a model or two; Korea has several AI companies with models of their own.

Cost matters now because there is a hundredfold gap between leading models and cheap ones. On the September 14 screen shown at the opening, the cost of solving the same task ranged from $0.06 to $8.75. What drives that cost is the number of output tokens spent finishing a task. Some models use 20 steps on a task and others use 100. AI labs may not particularly want you thinking this way, he added. They charge by the token.

You can no longer use price per token alone; it is outdated, as he put it. In a world of agents you have to think about output token volume and cost per task. Forces pushing cost down and up operate at the same time. Down: smaller models, sparsity, software efficiency. Up: larger models and growing reasoning tokens. One pattern was especially practical. Once a level of intelligence is reached, that intelligence gets cheap very quickly. Intelligence index 10 was reached in mid-to-late 2024, and cost has since fallen more than a hundredfold. If the top model feels too expensive now, you may not have to wait long.

The first question in the Q&A was the sharpest. Someone asked how performance is evaluated in the many languages that are not English. Ask for Korean expression, they said, and English idioms and English proverbs come out. George answered in two layers. He thinks these models abstract basic language understanding very well, because they abstract the concept of language itself in latent space. What he said next was candid. The nuance in how these models communicate is not captured well by the benchmarks that exist today, theirs included.

The second session was Chun Sung-jun of SK Telecom. The emphasis in his title sat on production-ready. This was not his first time here. At an NVIDIA Nemotron developer day in this same venue in April, he had said that the first cohort succeeded at building usable models and that the second cohort was about building models you can actually use. He had come back holding that promise.

A.X K2 has 688B parameters with 33B active and is published under Apache 2.0. The feedback they heard most after the first release was that the model was too big to use. That is really a worry about cost. Meanwhile the scale of the competition keeps growing. They could not shrink the model, so the answer had to come from architecture. Choose fewer tokens, store less state, control the outliers: three lines of attack. The result shows up directly in memory. The first cohort took 1038GB in BF16; this one takes 646GB at FP8 and 370GB at NVFP4, served on four B200s.

For usability, what they focused on was running as much improvement as possible inside a single training loop. Compute was not plentiful enough to run it several times. The one run had to work. The loop went from internal testing and Arena responses, to failure analysis mapping capability gaps, to updates through RLHF and RLVR, and back to evaluation. Observed behavior became the specification for the next training iteration. The result was a tie for first at 97.1 on AIME 2026, and long context, weak in the first cohort, rose 40 points.

The last session was Ahn Hong-jun of Trillion Labs, who opened by saying this was his first talk to a community this large. What they are building now is Gravity, a series aimed at AI factories. The future he is drawing is an AI agent that can genuinely self-drive complex physical infrastructure. They are deploying it at a 1.2-gigawatt thermal power plant. To define the problem they built an internal benchmark called PowerBench. The P&ID is the ground truth of that plant, and questions from the floor get answered on the drawing. The manuals and text do not hold the answer. If you cannot read the drawing, you cannot act in the plant.

Why the training pipeline has to run in that order was the clearest part of the night. The base VLM has never seen P&ID conventions. Reinforcement learning only sharpens what the policy can already sample, so rollouts from a policy that cannot read the drawing never reach the reward. You raise the prior first and run RL after. His doubt about designing a harness resolved once he saw the real data. The drawings are nearly A0 at 16 megapixels. Feeding one whole takes ages to prefill, and downscaling makes tags and symbols disappear. So they kept the overview and let the model zoom when it needs to.

For large-scale RL they went asynchronous. Agent rollouts are long-tailed, so every step waits on the slowest one. GPUs ran at around 30 percent. In the asynchronous setup workers stream trajectories without pause and the trainer does not wait. The key number is how many versions behind the trainer a trajectory may be. Most people set 1 or 2. They found a recipe that stretches to 8 and stayed stable. Throughput rose more than sixfold and GPU utilization went to 80 percent.

The last question reached the furthest point in the talk. Someone from a sound AI startup noted that in large power facilities, signs often appear in sound or vibration before telemetry catches them, and asked how that band would be handled. Calling it a genuinely good question, Ahn paused. Then he said that at this point he honestly does not know. Understanding it seems achievable, he thought, and finding an earlier indicator, predicting on it, and then acting is a very hard problem.

We closed with small group discussion. The loudest laugh came at how table leads were chosen. Whoever spends the most of their own money on AI leads the table, with company cards excluded. People stayed after the event ended and kept talking. It was a night with the team that builds the benchmark, the teams whose models appear on it, and the people who choose among those models all in one room.

Read the full write-up

Gallery

Price per token is the wrong unit now, photo 1Price per token is the wrong unit now, photo 2Price per token is the wrong unit now, photo 3Price per token is the wrong unit now, photo 4Price per token is the wrong unit now, photo 5Price per token is the wrong unit now, photo 6Price per token is the wrong unit now, photo 7Price per token is the wrong unit now, photo 8Price per token is the wrong unit now, photo 9Price per token is the wrong unit now, photo 10Price per token is the wrong unit now, photo 11Price per token is the wrong unit now, photo 12Price per token is the wrong unit now, photo 13Price per token is the wrong unit now, photo 14Price per token is the wrong unit now, photo 15Price per token is the wrong unit now, photo 16Price per token is the wrong unit now, photo 17Price per token is the wrong unit now, photo 18Price per token is the wrong unit now, photo 19Price per token is the wrong unit now, photo 20Price per token is the wrong unit now, photo 21Price per token is the wrong unit now, photo 22Price per token is the wrong unit now, photo 23Price per token is the wrong unit now, photo 24Price per token is the wrong unit now, photo 25Price per token is the wrong unit now, photo 26Price per token is the wrong unit now, photo 27Price per token is the wrong unit now, photo 28Price per token is the wrong unit now, photo 29Price per token is the wrong unit now, photo 30Price per token is the wrong unit now, photo 31Price per token is the wrong unit now, photo 32Price per token is the wrong unit now, photo 33Price per token is the wrong unit now, photo 34Price per token is the wrong unit now, photo 35Price per token is the wrong unit now, photo 36Price per token is the wrong unit now, photo 37Price per token is the wrong unit now, photo 38Price per token is the wrong unit now, photo 39Price per token is the wrong unit now, photo 40Price per token is the wrong unit now, photo 41Price per token is the wrong unit now, photo 42Price per token is the wrong unit now, photo 43Price per token is the wrong unit now, photo 44Price per token is the wrong unit now, photo 45Price per token is the wrong unit now, photo 46Price per token is the wrong unit now, photo 47Price per token is the wrong unit now, photo 48Price per token is the wrong unit now, photo 49Price per token is the wrong unit now, photo 50Price per token is the wrong unit now, photo 51Price per token is the wrong unit now, photo 52Price per token is the wrong unit now, photo 53Price per token is the wrong unit now, photo 54Price per token is the wrong unit now, photo 55Price per token is the wrong unit now, photo 56Price per token is the wrong unit now, photo 57Price per token is the wrong unit now, photo 58Price per token is the wrong unit now, photo 59Price per token is the wrong unit now, photo 60Price per token is the wrong unit now, photo 61Price per token is the wrong unit now, photo 62Price per token is the wrong unit now, photo 63Price per token is the wrong unit now, photo 64Price per token is the wrong unit now, photo 65Price per token is the wrong unit now, photo 66Price per token is the wrong unit now, photo 67Price per token is the wrong unit now, photo 68Price per token is the wrong unit now, photo 69Price per token is the wrong unit now, photo 70Price per token is the wrong unit now, photo 71

Next event