Back to the workflow
Quotaflow AI context
The company brain loaded into lemlist once, at step zero of the chain. Every prompt after it gets shorter because this document is already in the tool.
Quotaflow is a fictional company used as the running example for the webinar. This is what the output of step zero looks like in full, so you can see the shape before you point the prompt at your own website.
Company Overview
quotaflow.ai provides a continuous AI cost-optimization layer designed for companies shipping AI products and agents. The platform evaluates cheaper model paths, routes traffic to optimized providers, and measures realized savings without requiring code changes or new SDK integrations.
Positioning
The company positions itself as an essential infrastructure layer for "teams shipping AI products" that balances high quality with lower inference costs. It emphasizes a "quality-first" approach where cost savings are only implemented if they pass rigorous quality and performance benchmarks.
Core Offerings and Products/Services
- AI Gateway: A compatible API that routes approved traffic with built-in guardrails and fallback readiness.
- Experiment: A testing environment to compare baseline workloads against cheaper model candidates before production rollout.
- Procurement Agent: A tool to find better supply deals for the same model family currently in use.
- Usage Insights: A dashboard providing visibility into requests, tokens, spend, route mix, and realized savings.
- Shared-Savings Pilot: A low-risk entry point where the company charges based on verified savings.
Unique Value Proposition
- No Replatforming: Works with existing APIs, SDKs, and harnesses via a simple endpoint or key change.
- Verified Savings Loop: Only moves traffic to cheaper routes after they pass shadow/smoke testing for quality, latency, and success rate.
- Dual Optimization: Reduces costs by either finding better supply for the same model or a more efficient model mix for the same outcome.
- SLA-Backed: Provides enterprise-grade support and reliability for production workloads.
Target Market and Audience
- AI Product Companies: Teams needing to lower inference costs without changing product experience.
- AI Agent Companies: Developers requiring high reliability and quality while testing cheaper model paths.
- Agent Infrastructure Platforms: Providers looking to offer lower-cost routes to their own customers.
- Enterprise Buyers: Organizations requiring private gateways, compliant partners, and high-RPM failover.
Geography & Markets
- Home country: Undetermined (Uses .ai TLD and English language; no physical address or country-specific TLD provided).
- Primary markets served: Global/International, with specific model routing examples mentioning US, UK, DE (Germany), and CA (Canada).
Pains
- High Inference Costs: Fixes this by finding cheaper providers or more efficient model mixes (up to 50% savings).
- Integration Friction: Eliminates the need to rebuild apps or integrate new SDKs to achieve optimization.
- Quality Degradation Risk: Prevents quality loss by testing workloads against baselines before production changes.
- Lack of Visibility: Provides 100% usage visibility into tokens, spend, and "retry waste."
- Supply Volatility: Offers 24/7 supply visibility and managed fallback to ensure reliability.
Competitors
The market includes AI Gateway providers and LLM orchestration layers. While specific competitors aren't named, the company competes against internal engineering teams building their own optimization stacks and public API providers (OpenAI, Anthropic, Google) by offering discounted or optimized routes.
Income Sources
- Shared-Savings Model: Charges based on a percentage of verified savings realized by the customer.
- Enterprise Support: Likely includes fees for SLA-backed support and private gateway deployments.
- Partner Wallets: Settlement of qualified supply costs through the platform.
Additional Business Insights
- Security & Privacy: Features include encrypted keys (Server-side vault), no data retention (Enterprise safe mode), and no training on customer data.
- Model Agnostic: Supports leading labs including OpenAI, Anthropic, Google, Zhipu AI, ByteDance, Kimi, Kling, and Minimax.
Customer Stories & Social Proofs
- Tarris AI: "This is the first platform I’ve seen where we didn’t need to change any code or integrate another SDK to get evaluation, testing, and real cost savings—without sacrificing quality." — CEO, Tarris AI.
- Savings Metric: Claims up to "50% Inference cost optimization" and "38% for Claude family" in specific support agent use cases.
Site Structure
Product & Features
- /experiment — Experimentation and Testing
- /ai-procurement-agent — AI Procurement Agent
- /inference-supply-monitor — Supply Monitoring
- /route-audit — Route Auditing and Governance
- /team-context-sharing — Team Collaboration
Solutions & Use-Cases
- /case-studies — Workload Examples (Coding, Agents, Ops)
- /enterprise — Enterprise Solutions
- /partner — Partner and Wallet Solutions
- /connect-program — Connect Program
Technical & Models
Company & Pricing
- /pricing — Pricing and Savings Models
- /about — About the Company
- /contact-sales — Sales Contact
- /trust — Trust and Security Center
- /security — Security Protocols
GTM Motions
quotaflow sells cost savings that only exist once traffic is flowing. That shapes everything: the fastest path to revenue is getting a slice of real production traffic behind the gateway, not winning a long evaluation. Four motions follow from this.
Motion 1: Pilot-led outbound (primary XDR motion)
- Who: AI-native companies (Seed to Series C) already paying five to seven figures a month in inference, with agents or LLM features in production.
- Entry offer: the Shared-Savings Pilot. quotaflow only gets paid on verified savings, so the prospect's downside is close to zero.
- How it lands: one workload, shadow-tested against the current baseline, then routed to the cheaper path. Realized savings show up in Usage Insights within days.
- Sales cycle: two to six weeks from first reply to pilot live. Champion is technical, economic buyer is the CTO or finance.
- XDR job: find companies with a visible inference bill, reach the person who owns it, and book a 20-minute route audit or pilot scoping call.
Motion 2: Developer-led (self-serve pull)
- Who: individual engineers who discover the OpenAI-compatible endpoint through docs, GitHub, or community posts.
- Entry offer: swap the base URL and key, run Experiment against a real workload, see savings before committing.
- XDR job: monitor sign-ups and Experiment runs, qualify the account behind the sign-up, and convert high-spend accounts into a pilot or enterprise conversation. Treat every sign-up from a funded company as a warm lead.
Motion 3: Partner and channel
- Who: agent infrastructure platforms, AI dev tools, and agencies building agents for clients. They resell or embed cheaper routes for their own customers through the Partner Wallet.
- Entry offer: the Connect Program. One integration gives their whole customer base lower inference costs, and the partner earns on the spread or on retention.
- XDR job: identify platforms whose customers pay per token, and pitch margin protection for the platform itself, not just for end users.
Motion 4: Enterprise top-down
- Who: larger organizations with dedicated AI platform teams, compliance requirements, and multi-region traffic.
- Entry offer: private gateway, Enterprise safe mode (no data retention), compliant provider list, high-RPM failover, SLA.
- Sales cycle: two to six months, security review, procurement, buying committee.
- XDR job: multi-thread. Reach the AI platform lead, the FinOps or cloud cost owner, and the security or procurement function in parallel. Lead with reliability and governance (Route Audit), not with raw savings.
##
Where to spend your time
Motion 1 is where a new XDR should live. It produces pipeline fastest, the pilot removes the "prove it" objection, and every closed pilot becomes a case study for Motions 3 and 4\.
ICP Prioritization
| Tier | Company profile | Why they buy | Where to find them |
| Tier 1 | AI agent startups, Seed to Series B, 10 to 150 employees, agents in production, usage-based or seat pricing with heavy LLM usage per seat | Inference is their biggest COGS line and gross margin is what the next round is judged on | Recent funding announcements, "AI agent" in the tagline, YC and similar batch lists, model provider case studies |
| Tier 2 | Established SaaS (Series B to public) that added AI features in the last 18 months | AI features launched fast on the most expensive model; nobody has optimized yet, and finance is now asking questions | Product launch posts mentioning "powered by GPT/Claude/Gemini", hiring for "AI platform" or "LLM infra" roles |
| Tier 3 | Agent infra platforms, AI dev tools, agent agencies | Their customers' inference cost is their own margin problem; a cheaper route is a competitive feature | Directories of agent frameworks, no-code AI builders, dev tool marketplaces |
| Tier 4 | Enterprise with internal AI platform teams | Governance, failover, compliance, and cost visibility at scale | Job posts for "AI gateway", "LLM platform engineer", "GenAI CoE"; conference speaker lists |
Fast qualification rule: if you cannot find evidence that they run LLM calls in production at meaningful volume, they are not an account yet. Prototypes and internal chatbots do not have a bill worth optimizing.
Buyer Personas
Persona 1: The infra owner (champion, primary target)
- Titles: Head of AI Engineering, Staff or Lead ML Engineer, AI Platform Lead, Founding Engineer (at small startups), Head of Infrastructure.
- Reports to: CTO or VP Engineering.
- Decision role: champion. Runs the pilot, owns the routing decision, and often has the API keys.
- Personal pains:
- Gets pinged every time a provider has an outage or a rate-limit spike, and has no clean fallback.
- Knows the model mix is not optimal but has no time to build eval harnesses to prove a cheaper model is safe.
- Is asked by leadership "why is the OpenAI bill up 40 percent" and has no per-route breakdown to answer with.
- KPIs: uptime and latency of AI features, cost per request or per task, eval pass rates.
- What resonates: no code changes, shadow testing before any switch, per-route visibility, managed fallback.
- What kills it: anything that sounds like a rewrite, a new SDK, or a black box that swaps models without their approval.
- Best channel: email first (technical and to the point), LinkedIn second. They ignore fluff and respond to specificity about their stack.
Persona 2: The economic buyer (CTO or VP Engineering)
- Titles: CTO, VP Engineering, Co-founder and CTO.
- Decision role: signs off on the pilot and the enterprise agreement.
- Personal pains:
- Owns gross margin conversations with the CEO and board. Inference cost is often the line that decides whether the business model works.
- Wants optionality across model providers and does not want to be locked into one lab's pricing.
- Fears quality regressions that reach customers.
- KPIs: gross margin on AI features, engineering time spent on infra vs. product, incident count.
- What resonates: "pay only on verified savings", up to 50 percent optimization with quality gates, SLA, team time freed from building an internal router.
- Best channel: email with a one-line business case, or a warm intro from the champion.
Persona 3: The finance owner (CFO, Head of Finance, FinOps lead)
- Titles: CFO, VP Finance, Head of FinOps, Cloud Cost Manager.
- Decision role: influencer that can become a sponsor. Rarely the first contact, but a strong second thread in Tier 2 and Tier 4 accounts.
- Personal pains:
- AI spend is variable, hard to forecast, and shows up on a card or a single invoice with no unit economics behind it.
- Cannot tell if the AI team's spend is efficient or wasteful.
- KPIs: COGS, gross margin, forecast accuracy, cost per customer.
- What resonates: shared-savings pricing (no upfront cost), 100 percent visibility into tokens and spend, "retry waste" as a concrete leak.
- Best channel: email, short, numbers first. Never technical.
Persona 4: The founder-operator (Seed and Series A agent companies)
- Titles: CEO, Founder, Co-founder.
- Decision role: everything at once. At very small companies the founder is champion, buyer, and user.
- Personal pains: runway math. Every dollar saved on inference is a week of runway or a lower price point for customers.
- What resonates: the Tarris AI quote, fast time to value, no replatforming.
- Best channel: LinkedIn DM or short email; they read everything on their phone.
Persona priority for a new XDR: start with Persona 1 for the reply, loop in Persona 2 for the yes. Add Persona 3 as a second thread on Tier 2 and Tier 4 accounts.
Buying Triggers and Signals
Reach out when one of these is visible. Mention it in the first line.
| Signal | Why it matters | Where to spot it |
| New funding round (Seed to Series C) for an AI or agent company | Fresh scrutiny on burn and margin; hiring means more usage | Funding databases, press, LinkedIn announcements |
| Job post for "LLM infra", "AI platform engineer", "inference optimization", "evals" | They are about to build in-house what quotaflow already does | Job boards, company careers pages |
| Public post about token costs, rate limits, or provider outages | Pain is live and they are talking about it | LinkedIn, X, Hacker News, engineering blogs |
| Product launch of an agent or AI feature | Launched on the most expensive model; optimization has not happened yet | Product Hunt, launch posts, changelogs |
| Provider price change or outage in the news | Every team on that provider is reconsidering single-vendor risk that week | Provider status pages, tech press |
| Tech stack shows LiteLLM, OpenRouter, Helicone, Langfuse, or similar | They already believe in a gateway layer; the conversation is "why quotaflow" not "why a gateway" | Job posts, GitHub, docs, engineering blogs |
| Usage-based pricing on their own pricing page | Their margin depends directly on inference cost per unit | Their website |
| New CTO, VP Eng, or Head of AI hired | New leaders audit spend in their first 90 days | LinkedIn job changes |
Messaging by Persona
The core promise stays the same across personas: cut inference cost without touching code and without letting quality slip, and pay only when the savings are real. What changes is the angle and the proof.
Angle 1: The margin angle (CTO, CFO, Founder)
- Core tension: the AI feature is popular, and every new user makes gross margin worse.
- Hook: "Most agent companies we talk to run 100 percent of traffic on the model they prototyped with. On support-agent workloads that gap has been worth up to 38 percent on the Claude family alone."
- CTA: "Want a route audit on one workload? We only bill on savings we can show you in your own dashboard."
Angle 2: The "you're about to build it" angle (Infra owner, CTO)
- Core tension: they are hiring or assigning engineers to build routing, evals, and fallback themselves.
- Hook: "Saw you're hiring an AI platform engineer. Curious whether the routing and eval layer is on that roadmap, or whether you'd rather that person ship product."
- CTA: "Happy to show how a team like \[peer\] got shadow-tested routing live in an afternoon with a key swap."
Angle 3: The reliability angle (Infra owner, Enterprise)
- Core tension: one provider, one region, no fallback, and the on-call rotation feels it.
- Hook: "When \[provider\] degraded last week, did your agents fail over or did they just fail?"
- CTA: "We can put managed fallback and 24/7 supply visibility in front of one workload without changing a line of code. Worth 20 minutes?"
Angle 4: The visibility angle (Finance, Enterprise)
- Core tension: AI spend is a single line on an invoice with no unit economics behind it.
- Hook: "Can you see today what a single customer conversation costs you in tokens, including retries? Most finance teams we meet cannot."
- CTA: "We give 100 percent usage visibility as part of the pilot. No cost until savings are verified."
Angle 5: The single-vendor risk angle (CTO, Founder)
- Core tension: locked into one lab's pricing and roadmap.
- Hook: "You're one pricing change away from your unit economics breaking. quotaflow keeps your current model where it earns its place and moves the rest to verified cheaper routes."
Copy rules for this product
- Always say "no code changes" and "shadow-tested before any switch" early. Those two phrases dissolve the two biggest fears.
- Never promise 50 percent. Say "up to 50 percent" and "38 percent on the Claude family in support-agent use cases". Be specific and honest.
- The Shared-Savings Pilot is the CTA of choice. It converts skepticism into a low-stakes yes.
- Lead with their workload, not our product list. "Your support agent" beats "our AI Gateway".
Discovery Questions
Use these on the first call or in a follow-up email. The goal is to size the bill and find the workload for the pilot.
1. Which models and providers are you running in production today, and roughly what does that cost per month?
2. Which workload eats the biggest share of that spend? (support agent, coding agent, RAG, ops automation)
3. Have you tested a cheaper model on that workload? What stopped you from switching?
4. How do you evaluate quality today when you change a prompt or a model?
5. What happens when your provider rate-limits you or goes down?
6. Can you see cost per request, per customer, or per task right now?
7. Who else cares about this number? (surfaces the CTO or CFO for multi-threading)
8. Are you on a committed-spend agreement with any provider? When does it renew?
Pilot fit checklist: a real production workload, at least one identifiable owner, a measurable baseline, and no hard contractual lock-in on 100 percent of traffic.
Objection Handling
"We built our own router / we use LiteLLM." Great, so you already believe in the layer. The question is who maintains the eval harness, the provider deals, and the fallback logic. quotaflow sits on the same OpenAI-compatible interface, so you can put one workload behind it and compare. If your in-house setup wins, you've lost an afternoon.
"We can't risk quality on customer-facing agents." Neither can we, which is why nothing moves until it passes shadow and smoke tests against your baseline on quality, latency, and success rate. You approve the route. Quality-first is the product's whole premise.
"Our data can't leave our environment / security will block this." Enterprise safe mode means no data retention, keys stay in a server-side vault, and we never train on customer data. Private gateway deployments exist for exactly this. Let's get security in the room early.
"We have a committed-spend deal with OpenAI / Anthropic." Then the Procurement Agent angle applies: same model family, better supply terms. And committed spend rarely covers 100 percent of traffic. We optimize what sits outside the commitment and help you size the next renewal with real data.
"Our AI bill isn't big enough to matter." It will be, and the cheapest time to put the layer in is before the bill is scary. That said, if you are under a few thousand a month, be honest, park the account, and set a signal on funding or hiring.
"Adding a gateway adds latency." Latency is one of the three gates a route has to pass before it goes live. If a route is slower than your baseline it does not get approved. Ask for their current p95 and offer to measure it in Experiment.
"We'll just switch models ourselves." Ask how they will validate that switch, and what it costs in engineering time each time a new model ships. quotaflow does that loop continuously, and the model landscape changes monthly.
"Not a priority right now." Agree, then anchor: "Understood. Most teams pick this up right after a funding round or a margin review. Mind if I check back in then? In the meantime, here is a two-minute route audit on your public pricing model." Leave value, keep the thread.
Competitor Battlecards
quotaflow competes across three tiers. The most common competitor is not a vendor, it is the engineer who thinks they can build it in a weekend.
Competitive Set Summary
| Player | Tier | Their positioning | Primary ICP | quotaflow's advantage |
| OpenRouter | Direct | Unified API to hundreds of models, pay-as-you-go | Indie devs, startups wanting breadth | Verified savings loop and quality gates; OpenRouter routes, it does not test |
| Portkey | Direct | AI gateway with observability, guardrails, governance | Mid-market and enterprise platform teams | Shared-savings pricing and active cost optimization; Portkey charges a platform fee regardless of savings |
| LiteLLM (open source) | Direct / Indirect | Free, self-hosted proxy for 100+ providers | Engineering teams that want control | Zero maintenance, procurement deals, SLA; LiteLLM is a build project you own forever |
| Helicone / Langfuse | Indirect | Observability and logging for LLM apps | Teams debugging quality and cost | They show you the bill; quotaflow lowers it |
| Cloudflare / Vercel AI Gateway | Indirect | Gateway bundled with a hosting or edge platform | Teams already on that platform | Model-agnostic, provider-agnostic, and optimization is the product, not a feature |
| Martian / Not Diamond | Direct | Intelligent model routing per request | Teams optimizing cost-quality trade-offs | Broader loop (procurement, supply monitoring, fallback, audit) plus pay-on-savings |
| Hyperscalers (Bedrock, Vertex, Azure OpenAI) | Indirect | Managed model access inside your cloud | Enterprises with cloud commitments | Cross-provider optimization; hyperscalers keep you inside their catalog |
| In-house build | Indirect | "Our engineers can handle it" | Every well-funded startup | Time to value in an afternoon vs. months, continuous eval loop, no headcount |
| Status quo | Status quo | One model, one provider, no testing | Everyone else | Up to 50 percent savings sitting on the table with zero code changes |
Battlecard: OpenRouter
- Strengths (be honest): massive model catalog, very fast to start, strong developer mindshare, transparent per-model pricing.
- Weaknesses: no quality gating before switching models, no shadow testing, no procurement negotiation, limited enterprise controls.
- Why prospects pick them: breadth and simplicity.
- Why they pick quotaflow instead: they are past the experimentation phase and need savings that are verified, governed, and safe for production.
- Outreach hook: "OpenRouter gets you access to every model. Have you measured which of them you should actually be on for your support agent, and what it costs you not to know?"
Battlecard: Portkey
- Strengths: mature observability, guardrails, prompt management, enterprise governance features, solid docs.
- Weaknesses: priced as a platform subscription independent of outcome; optimization is a feature set you configure, not a managed loop.
- Why prospects pick them: they want one control plane for everything LLM.
- Why they pick quotaflow instead: they want the cost to go down, not another dashboard to configure, and they like paying only on results.
- Outreach hook: "Portkey shows you where the spend goes. Who on your team is turning that into a cheaper route mix every month?"
Battlecard: LiteLLM and in-house builds
- Strengths: free, flexible, full control, active community.
- Weaknesses: someone owns the upgrades, the eval harness, the provider relationships, and the 3 a.m. fallback logic. Savings depend entirely on your team's time.
- Why prospects pick them: engineering pride and perceived zero cost.
- Why they pick quotaflow instead: the engineer they'd assign is worth more shipping product; quotaflow is OpenAI-compatible, so migration is a key swap.
- Outreach hook: "Running LiteLLM? Then you're one config away from testing whether a managed loop beats what your team maintains. We only bill on the difference."
Battlecard: Observability tools (Helicone, Langfuse)
- Positioning: not a replacement; a signal. Teams using them care about cost and quality and already log everything.
- Outreach hook: "You can see exactly what each request costs. Want the layer that acts on it?"
Battlecard: Hyperscaler gateways
- Strengths: compliance, procurement simplicity, cloud credits.
- Weaknesses: catalog limited to the models the cloud sells; little incentive to route you to a cheaper competitor.
- Outreach hook: "Bedrock and Vertex will never tell you a cheaper route exists outside their catalog. We will, and we test it first."
Where quotaflow loses (know this before you pitch)
- Very early teams with negligible spend: LiteLLM or OpenRouter is the right call. Park them.
- Teams whose main pain is prompt management and tracing, not cost: Portkey or Langfuse wins today. Keep the thread for when the bill grows.
- Deeply committed hyperscaler accounts with 100 percent traffic under contract until renewal: time the outreach to the renewal.
Battlecard summary for the first 30 seconds
- Lead with: verified savings loop (shadow test before switch) and pay-only-on-savings.
- Avoid leading with: "we support many models" or "we have a dashboard". Every gateway says that.
- Proof: Tarris AI quote (no code changes, no new SDK), up to 50 percent optimization, 38 percent on the Claude family in support-agent workloads.
#
Cold Call Talk Track
Opener (permission-based): "Hi \[Name\], it's \[You\] from quotaflow. I'll be quick: I work with AI teams like \[peer company\] on inference cost. Do you have 30 seconds for me to say why I called, and you tell me if it's relevant?"
Reason for the call (pick the matching trigger):
- Funding: "Congrats on the round. Usually right after, the board starts asking about gross margin on the AI side. Is that on your desk?"
- Hiring: "Saw you're hiring for AI platform. Is model routing and evals part of what that person will build?"
- Outage: "When \[provider\] had issues last week, did your agents fail over or just fail?"
Pain probe: "Roughly what share of your traffic runs on the model you started with? And has anyone tested a cheaper path on your biggest workload?"
Mini proof: "We put a testing and routing layer in front of your current API with a key swap. Nothing switches until it passes quality, latency, and success rate against your baseline. On support agents we've seen up to 38 percent on the Claude family."
Ask: "The way we start is a shared-savings pilot on one workload. You pay only on savings you can see in your own dashboard. Would a 20-minute scoping call this week or next make sense?"
Qualification and Handoff
Qualified for a pilot (pass to AE) when:
- LLM traffic in production with an identifiable workload.
- Spend is meaningful (rule of thumb: several thousand USD a month or growing fast).
- A technical owner is engaged and willing to run Experiment.
- No 100 percent lock-in on provider contracts.
Disqualify or nurture when:
- Prototype stage, no production traffic.
- Bill is trivial and no growth signal.
- Hard regulatory constraint that even private gateway and safe mode cannot satisfy (rare; confirm with the team first).
Handoff notes should include: models and providers in use, estimated monthly spend, the candidate pilot workload, current gateway or observability tooling, contract renewal dates, and everyone who touched the thread.
#
Glossary for XDR Newcomers
- Inference: running a model to get an output. Every request to GPT, Claude, or Gemini is inference, and it is what the customer pays for.
- Token: the unit models bill on. Roughly three quarters of a word. Input tokens (what you send) and output tokens (what comes back) are priced differently.
- Model family / model mix: a family is one lab's lineup (the Claude family, the GPT family). Model mix is which models handle which parts of a workload.
- Provider / supply: who serves the model. The same model can be available from the lab directly, a hyperscaler, or a specialized inference provider at different prices and speeds.
- AI Gateway: a proxy that sits between the customer's app and the model providers. All requests go through it, so it can route, log, retry, and fail over.
- OpenAI-compatible: the gateway speaks the same API format as OpenAI, so almost any app can point at it by changing a URL and a key. This is why "no code changes" is true.
- Routing: choosing which model or provider handles a request.
- Shadow testing: sending a copy of real traffic to a candidate route and comparing results without the customer ever seeing the candidate's output.
- Smoke test: a quick sanity check that a route works before broader rollout.
- Fallback / failover: automatically switching to another provider when the primary one errors, slows down, or rate-limits.
- RPM / rate limit: requests per minute a provider allows. High-RPM failover matters for agents that fire many calls at once.
- Latency / p95: how long a request takes. p95 means 95 percent of requests are faster than this number. Agents are latency-sensitive.
- Retry waste: tokens paid for on requests that failed and were retried. Invisible on a normal invoice.
- Eval / benchmark: a test set used to judge whether a model's outputs are good enough. The "quality gate" in quotaflow's pitch.
- COGS / gross margin: cost of serving a customer, and revenue minus that cost. Inference is often the biggest COGS line for AI companies, which is why CFOs care.
- Committed spend: a contract where a company prepays or commits to a volume with a provider for a discount. It creates lock-in and renewal-date timing opportunities.
- Shared savings: quotaflow's pricing model. Fees are a percentage of savings that have been verified in the dashboard. No savings, no fee.
- Agent: an AI system that takes multiple steps and calls tools to complete a task. Agents make many model calls per task, so their bills add up fast.