AI-mallit ja alustat
Puuttuva mittari tokenien ja pilvikulujen välillä

Ongelma ei ole, että tekoälytiimeillä ei ole kustannustietoja. Ongelma on, että tokenien hallintapaneeli ja pilvilasku kuvaavat eri järjestelmiä, jotka kuuluvat eri tiimeille, eikä niillä ole luotettavaa tapaa yhdistää niitä.
Tukihenkilö voi ratkaista yhden tiketin viiden mallikutsun, yhden haun, kahden työkalukutsun ja yhden uudelleenyrityksen jälkeen. Liiketoiminta kirjaa yhden valmiin tapauksen. Infrastruktuuri kirjaa hajanaisia pyyntöjä, podseja, muistia, kiihdytysaikaa ja jaettuja palveluita. Kun nämä tiedot eivät kohtaa, kustannusoptimointi on osittain arvailua.
Miksi tokenimittarit ja pilvilaskut kertovat eri tarinoita?
Tokenimäärät ovat hyödyllisiä. Ne näyttävät, kuinka paljon tekstiä malli vastaanotti ja palautti, ja ne auttavat tiimejä vertailemaan kehotteita, malleja tai reititysvaihtoehtoja. Mutta ne eivät kerro, mitä tapahtui mallikutsun ympärillä, kuinka paljon laskentatehoa tuki hakua ja työkalujen käyttöä, kuinka monta epäonnistunutta yritystä tapahtui ensin, tai oliko lopputulos hyödyllinen.
The State of FinOps 2026 osoittaa, kuinka nopeasti tekoäly on siirtynyt tavalliseen FinOps-työhön: 98 % vastaajista hallitsee nyt tekoälykuluja, verrattuna 63 %:iin vuonna 2025. Mutta suurempi budjettirivi ei silti kerro, mikä työnkulku poltti rahaa tai miksi.
Kaksi asiakirjankäsittelytyötä voi käyttää suunnilleen saman määrän tokeneita. Yksi voi päättyä yhteen mallipyyntöön. Toinen voi hakea kontekstia useista varastoista, kutsua ulkoista palvelua, siirtyä toiseen malliin ja suorittaa asiakirjan uudelleen epäonnistuneen validointitarkistuksen jälkeen, jonka käyttäjä ei koskaan näe. Tokenien kokonaismäärät näyttävät samankaltaisilta, mutta toteutuspolut eivät ole.
Unite.ai on jo tutkinut, miksi tokenimäärät eivät automaattisesti edusta liiketoiminnan arvoa. Seuraava vaihe on yhdistää nämä määrät niiden tuottaneisiin työkuormiin. Muuten tiimi voi parantaa kustannusta per token, mutta heikentää valmiin tehtävän kustannusta.
Miltä täydellinen kustannusketju näyttää?
Hyödyllinen kustannusketju alkaa liiketoimintaa kiinnostavasta tuloksesta. Se voi olla ratkaistu tukipyyntö, käsitelty asiakirja, hyväksytty koodimuutos tai valmis agenttityönkulku. Kaikki sen alapuolella tarvitsee tunnisteen, jota voidaan seurata järjestelmän läpi.
Sovelluskerros tarjoaa ensimmäisen yhteyden. Pyyntö‑tunnus, jäljitystunnus, työnkulun nimi tai keskustelutunnus voivat yhdistää useita mallin ja työkalun operaatioita yhteen työyksikköön. Ilman tätä lankaa kymmenen toisiinsa liittyvää tapahtumaa näyttää kymmeneltä erilliseltä maksulta.
The OpenTelemetry‑konventiot GenAI‑agenteille tarjoavat nousevan sanaston tälle kerrokselle. Ne kattavat operaatioita, tarjoajia, pyydettyjä malleja, agenteja, keskusteluja, token‑käyttöä, työkalujen suorittamista, virheitä ja työnkulkuja. Konventiot ovat edelleen kehitysvaiheessa, joten tiimien ei tulisi pitää niitä valmiina universaalina standardina. Ne ovat hyödyllisiä, koska ne konkretisoivat korrelaatio‑ongelman.
Then comes infrastructure. AWS’s split cost allocation data for EKS voi kohdistaa jaetut laskenta‑ ja muistikulut Kubernetes‑podeille ja paljastaa tietoja, kuten klusteri, nimiavaruus, käyttöönotto, solmu, työkuorman nimi ja tyyppi. Tuetuille kiihdytyksellisille instansseille tiedot kattavat myös GPU‑, Trainium‑ ja Inferentia‑varaukset.
That’s the other half of the chain. A trace can explain what the application tried to do; Kubernetes allocation can show which resources carried the work. Unite.ai’s guide to deploying and monitoring LLMs on Kubernetes provides the wider production context, including resource allocation, scaling, and observability.
The join won’t happen by accident. Teams need a stable identifier that survives long enough to connect application telemetry with workload labels, allocation records, or another mapping layer. Customer data doesn’t belong in Kubernetes tags. Teams should decide which low-cardinality identifiers can safely connect a workflow category, service, or feature to the resources it consumed.
Once that application context is in place, teams can start tracking Kubernetes costs by workload and connect namespace, CPU, memory, and GPU usage back to the work being performed. That still doesn’t tell you whether the workflow created business value, but it gives the infrastructure side of the calculation something concrete to attach to.
Mihin yksikkömittariin liiketoiminnan tulisi luottaa?
Ei ole yhtä ainoaa tekoälykustannusmittaria, jota jokaisen tiimin tulisi käyttää. Kustannus per token vastaa mallin kulutuskysymykseen. Kustannus per pod vastaa infrastruktuurin allokointikysymykseen. Kumpikaan ei kerro tuoteomistajalle, tuottaako ominaisuus riittävän arvon.
The best denominator is usually the smallest outcome the business can define clearly, and the product team can influence. A support operation might track cost per resolved case. A document system might use cost per successfully processed file, while a coding assistant could examine cost per accepted change rather than cost per suggestion.
Success changes the math.
A workflow with a low cost per attempt may be expensive if it fails often, triggers repeated validation, or sends too many cases to human review. That is why teams should separate cost to attempt from cost to complete and, where possible, cost per accepted outcome. The last number is often the most useful because it includes the work the system produced but the business couldn’t use.
Agent systems make this harder because their paths can change from one run to the next. Unite.ai’s analysis of the economics of scaling agentic AI workloads covers routing, tool calls, retries, and workflow-level attribution. Those behaviors belong in the unit metric when they consume resources, even when the final user sees only one answer.
The metric still won’t be perfect. Shared services, cached results, batch jobs, and delayed processing can blur attribution. A decision-useful estimate is better than false precision, especially when it tells engineers which layer deserves investigation.
Kuka omistaa luvun?
The hardest part may be organizational. ML teams understand model calls and evaluation. Platform teams understand workloads and cluster behavior. FinOps understands billing data and allocation rules. Product teams know which outcome matters.
No one team owns the full chain.
That creates a predictable argument over whose dashboard is correct. The ML team may point to lower token use, while the platform team sees GPU hours climbing and the product team sees fewer completed tasks than before. All three observations can be true at once. The shared metric has to explain the relationship between them.
A workable starting point is one production workflow with a clear completion event. Give it a stable identifier. Carry that context through the model and tool traces, map it to the service or workload running in Kubernetes, and choose one business denominator. Then bring the teams together when the number moves unexpectedly.
That review matters more than a polished dashboard. A sudden increase may come from longer prompts, a new fallback path, underused GPU capacity, a changed autoscaling policy, or a product decision that sends more work through the AI feature. Each cause belongs to a different owner.
Automation should come later. A recommendation engine can only act on the labels and thresholds it receives, and a bad denominator can make an efficient system look wasteful or reward a cheap workflow that users reject. Teams need enough shared visibility to distinguish model behavior from application design and infrastructure allocation before they let a system act on the result. Otherwise, an automated cost fix can reduce capacity, raise latency, and move the expense somewhere less visible.
Kustannusketjun on oltava jaettu
AI cost control will stay fragmented as long as every team optimizes only the layer it can see. Tokens, traces, pods, accelerators, and invoices aren’t rival measurements. They are pieces of the same cost chain.
The companies that connect them won’t get a perfect number on day one. What matters is whether the team can trace a high bill back to the workflow that caused it, work out what changed, and decide if the result justified the cost.












