Back to the workflow

Quotaflow AI context

The company brain loaded into lemlist once, at step zero of the chain. Every prompt after it gets shorter because this document is already in the tool.

Quotaflow is a fictional company used as the running example for the webinar. This is what the output of step zero looks like in full, so you can see the shape before you point the prompt at your own website.

Company Overview

quotaflow.ai provides a continuous AI cost-optimization layer designed for companies shipping AI products and agents. The platform evaluates cheaper model paths, routes traffic to optimized providers, and measures realized savings without requiring code changes or new SDK integrations.

Positioning

The company positions itself as an essential infrastructure layer for "teams shipping AI products" that balances high quality with lower inference costs. It emphasizes a "quality-first" approach where cost savings are only implemented if they pass rigorous quality and performance benchmarks.

Core Offerings and Products/Services

Unique Value Proposition

Target Market and Audience

Geography & Markets

Pains

Competitors

The market includes AI Gateway providers and LLM orchestration layers. While specific competitors aren't named, the company competes against internal engineering teams building their own optimization stacks and public API providers (OpenAI, Anthropic, Google) by offering discounted or optimized routes.

Income Sources

Additional Business Insights

Customer Stories & Social Proofs

Site Structure

Product & Features

Solutions & Use-Cases

Technical & Models

Company & Pricing

GTM Motions

quotaflow sells cost savings that only exist once traffic is flowing. That shapes everything: the fastest path to revenue is getting a slice of real production traffic behind the gateway, not winning a long evaluation. Four motions follow from this.

Motion 1: Pilot-led outbound (primary XDR motion)

Motion 2: Developer-led (self-serve pull)

Motion 3: Partner and channel

Motion 4: Enterprise top-down

##

Where to spend your time

Motion 1 is where a new XDR should live. It produces pipeline fastest, the pilot removes the "prove it" objection, and every closed pilot becomes a case study for Motions 3 and 4\.

ICP Prioritization

TierCompany profileWhy they buyWhere to find them
Tier 1AI agent startups, Seed to Series B, 10 to 150 employees, agents in production, usage-based or seat pricing with heavy LLM usage per seatInference is their biggest COGS line and gross margin is what the next round is judged onRecent funding announcements, "AI agent" in the tagline, YC and similar batch lists, model provider case studies
Tier 2Established SaaS (Series B to public) that added AI features in the last 18 monthsAI features launched fast on the most expensive model; nobody has optimized yet, and finance is now asking questionsProduct launch posts mentioning "powered by GPT/Claude/Gemini", hiring for "AI platform" or "LLM infra" roles
Tier 3Agent infra platforms, AI dev tools, agent agenciesTheir customers' inference cost is their own margin problem; a cheaper route is a competitive featureDirectories of agent frameworks, no-code AI builders, dev tool marketplaces
Tier 4Enterprise with internal AI platform teamsGovernance, failover, compliance, and cost visibility at scaleJob posts for "AI gateway", "LLM platform engineer", "GenAI CoE"; conference speaker lists

Fast qualification rule: if you cannot find evidence that they run LLM calls in production at meaningful volume, they are not an account yet. Prototypes and internal chatbots do not have a bill worth optimizing.

Buyer Personas

Persona 1: The infra owner (champion, primary target)

Persona 2: The economic buyer (CTO or VP Engineering)

Persona 3: The finance owner (CFO, Head of Finance, FinOps lead)

Persona 4: The founder-operator (Seed and Series A agent companies)

Persona priority for a new XDR: start with Persona 1 for the reply, loop in Persona 2 for the yes. Add Persona 3 as a second thread on Tier 2 and Tier 4 accounts.

Buying Triggers and Signals

Reach out when one of these is visible. Mention it in the first line.

SignalWhy it mattersWhere to spot it
New funding round (Seed to Series C) for an AI or agent companyFresh scrutiny on burn and margin; hiring means more usageFunding databases, press, LinkedIn announcements
Job post for "LLM infra", "AI platform engineer", "inference optimization", "evals"They are about to build in-house what quotaflow already doesJob boards, company careers pages
Public post about token costs, rate limits, or provider outagesPain is live and they are talking about itLinkedIn, X, Hacker News, engineering blogs
Product launch of an agent or AI featureLaunched on the most expensive model; optimization has not happened yetProduct Hunt, launch posts, changelogs
Provider price change or outage in the newsEvery team on that provider is reconsidering single-vendor risk that weekProvider status pages, tech press
Tech stack shows LiteLLM, OpenRouter, Helicone, Langfuse, or similarThey already believe in a gateway layer; the conversation is "why quotaflow" not "why a gateway"Job posts, GitHub, docs, engineering blogs
Usage-based pricing on their own pricing pageTheir margin depends directly on inference cost per unitTheir website
New CTO, VP Eng, or Head of AI hiredNew leaders audit spend in their first 90 daysLinkedIn job changes

Messaging by Persona

The core promise stays the same across personas: cut inference cost without touching code and without letting quality slip, and pay only when the savings are real. What changes is the angle and the proof.

Angle 1: The margin angle (CTO, CFO, Founder)

Angle 2: The "you're about to build it" angle (Infra owner, CTO)

Angle 3: The reliability angle (Infra owner, Enterprise)

Angle 4: The visibility angle (Finance, Enterprise)

Angle 5: The single-vendor risk angle (CTO, Founder)

Copy rules for this product

Discovery Questions

Use these on the first call or in a follow-up email. The goal is to size the bill and find the workload for the pilot.

1. Which models and providers are you running in production today, and roughly what does that cost per month?

2. Which workload eats the biggest share of that spend? (support agent, coding agent, RAG, ops automation)

3. Have you tested a cheaper model on that workload? What stopped you from switching?

4. How do you evaluate quality today when you change a prompt or a model?

5. What happens when your provider rate-limits you or goes down?

6. Can you see cost per request, per customer, or per task right now?

7. Who else cares about this number? (surfaces the CTO or CFO for multi-threading)

8. Are you on a committed-spend agreement with any provider? When does it renew?

Pilot fit checklist: a real production workload, at least one identifiable owner, a measurable baseline, and no hard contractual lock-in on 100 percent of traffic.

Objection Handling

"We built our own router / we use LiteLLM." Great, so you already believe in the layer. The question is who maintains the eval harness, the provider deals, and the fallback logic. quotaflow sits on the same OpenAI-compatible interface, so you can put one workload behind it and compare. If your in-house setup wins, you've lost an afternoon.

"We can't risk quality on customer-facing agents." Neither can we, which is why nothing moves until it passes shadow and smoke tests against your baseline on quality, latency, and success rate. You approve the route. Quality-first is the product's whole premise.

"Our data can't leave our environment / security will block this." Enterprise safe mode means no data retention, keys stay in a server-side vault, and we never train on customer data. Private gateway deployments exist for exactly this. Let's get security in the room early.

"We have a committed-spend deal with OpenAI / Anthropic." Then the Procurement Agent angle applies: same model family, better supply terms. And committed spend rarely covers 100 percent of traffic. We optimize what sits outside the commitment and help you size the next renewal with real data.

"Our AI bill isn't big enough to matter." It will be, and the cheapest time to put the layer in is before the bill is scary. That said, if you are under a few thousand a month, be honest, park the account, and set a signal on funding or hiring.

"Adding a gateway adds latency." Latency is one of the three gates a route has to pass before it goes live. If a route is slower than your baseline it does not get approved. Ask for their current p95 and offer to measure it in Experiment.

"We'll just switch models ourselves." Ask how they will validate that switch, and what it costs in engineering time each time a new model ships. quotaflow does that loop continuously, and the model landscape changes monthly.

"Not a priority right now." Agree, then anchor: "Understood. Most teams pick this up right after a funding round or a margin review. Mind if I check back in then? In the meantime, here is a two-minute route audit on your public pricing model." Leave value, keep the thread.

Competitor Battlecards

quotaflow competes across three tiers. The most common competitor is not a vendor, it is the engineer who thinks they can build it in a weekend.

Competitive Set Summary

PlayerTierTheir positioningPrimary ICPquotaflow's advantage
OpenRouterDirectUnified API to hundreds of models, pay-as-you-goIndie devs, startups wanting breadthVerified savings loop and quality gates; OpenRouter routes, it does not test
PortkeyDirectAI gateway with observability, guardrails, governanceMid-market and enterprise platform teamsShared-savings pricing and active cost optimization; Portkey charges a platform fee regardless of savings
LiteLLM (open source)Direct / IndirectFree, self-hosted proxy for 100+ providersEngineering teams that want controlZero maintenance, procurement deals, SLA; LiteLLM is a build project you own forever
Helicone / LangfuseIndirectObservability and logging for LLM appsTeams debugging quality and costThey show you the bill; quotaflow lowers it
Cloudflare / Vercel AI GatewayIndirectGateway bundled with a hosting or edge platformTeams already on that platformModel-agnostic, provider-agnostic, and optimization is the product, not a feature
Martian / Not DiamondDirectIntelligent model routing per requestTeams optimizing cost-quality trade-offsBroader loop (procurement, supply monitoring, fallback, audit) plus pay-on-savings
Hyperscalers (Bedrock, Vertex, Azure OpenAI)IndirectManaged model access inside your cloudEnterprises with cloud commitmentsCross-provider optimization; hyperscalers keep you inside their catalog
In-house buildIndirect"Our engineers can handle it"Every well-funded startupTime to value in an afternoon vs. months, continuous eval loop, no headcount
Status quoStatus quoOne model, one provider, no testingEveryone elseUp to 50 percent savings sitting on the table with zero code changes

Battlecard: OpenRouter

Battlecard: Portkey

Battlecard: LiteLLM and in-house builds

Battlecard: Observability tools (Helicone, Langfuse)

Battlecard: Hyperscaler gateways

Where quotaflow loses (know this before you pitch)

Battlecard summary for the first 30 seconds

#

Cold Call Talk Track

Opener (permission-based): "Hi \[Name\], it's \[You\] from quotaflow. I'll be quick: I work with AI teams like \[peer company\] on inference cost. Do you have 30 seconds for me to say why I called, and you tell me if it's relevant?"

Reason for the call (pick the matching trigger):

Pain probe: "Roughly what share of your traffic runs on the model you started with? And has anyone tested a cheaper path on your biggest workload?"

Mini proof: "We put a testing and routing layer in front of your current API with a key swap. Nothing switches until it passes quality, latency, and success rate against your baseline. On support agents we've seen up to 38 percent on the Claude family."

Ask: "The way we start is a shared-savings pilot on one workload. You pay only on savings you can see in your own dashboard. Would a 20-minute scoping call this week or next make sense?"

Qualification and Handoff

Qualified for a pilot (pass to AE) when:

Disqualify or nurture when:

Handoff notes should include: models and providers in use, estimated monthly spend, the candidate pilot workload, current gateway or observability tooling, contract renewal dates, and everyone who touched the thread.

#

Glossary for XDR Newcomers