Cost Economics · 2 min read

AI inference is COGS. Start treating it like one.

Most AI-native teams bury inference spend in "R&D" or a cloud line item. But it's a cost of goods sold — and where you book it changes whether you actually know your margins.

By Akhil Anand · August 21, 2026

Ask a traditional SaaS founder what their gross margin is and they'll answer in seconds. Ask an AI-native founder the same thing and you'll often get a pause — because the single biggest variable cost of serving a customer is sitting in the wrong bucket.

Where the cost hides

When you serve a request, the model call is a direct cost of delivering the product. That's the textbook definition of cost of goods sold. Yet inference spend routinely gets filed under "R&D," "infrastructure," or a generic cloud line — anywhere but COGS.

It feels harmless. It isn't. The moment inference leaves COGS, your gross margin stops describing reality. You report a healthy margin that quietly excludes the fastest-growing cost you have.

A gross margin that doesn't include inference isn't conservative. It's fiction — and it's the number your pricing decisions rest on.

Per-request COGS is the unit that matters

Booking inference as COGS forces a useful question: what does one unit of your product cost to serve? Not per month, not per team — per request, per session, per user action. Once you have that, everything downstream sharpens:

  • Pricing: you can tell which plans are profitable and which quietly subsidize power users.
  • Margin tracking: you see erosion as it happens, not two quarters later.
  • Feature economics: you learn which features carry their weight and which are loss leaders you never chose.

Margin erosion is a COGS problem

Here's why the classification isn't just accounting hygiene. AI COGS moves — with model mix, context growth, and usage depth — in a way rent and salaries don't. If inference is buried in opex, rising costs look like "we're spending more on infra." If it's in COGS, the same rise shows up as margin compression, which is what it actually is, and which is the thing that ends companies.

Treat inference as COGS, measure it per request, and margin stops being a vibe and becomes a number you can defend. That's the whole point of unit economics — and for AI-native companies, inference is the unit.


Try AtlasBurn free