All notes

#207 Building vertical AI & AI-native services

September 1, 2026·5 min read

#207 — Building vertical AI & AI-native services

Vertical AI collapses software and services into a unified delivery layer, expanding addressable markets from constrained 1% to 3% IT software budgets into 20% to 50%+ operational labor line items. Founders can capture these budgets either by licensing vertical AI workflow tools to customer teams or by operating as a full-stack, AI-native service provider that guarantees the end outcome.

The Macro Shift

Legacy SaaS digitized back-office systems of record, while vertical AI automates domain-specific knowledge work directly.

  • Market expansion: U.S. professional and business services account for 13% of GDPa market roughly 10x larger than the global software sector.
  • High-velocity scaling: Breakout category leaders reach $10M to $100M+ ARR substantially faster than traditional SaaS by pricing directly against work output.
  • Prime target verticals: Regulated, document-dense sectors with repetitive heuristicssuch as insurance brokerage, legal discovery, fund administration, revenue cycle management, and construction estimationpresent immediate disruption potential.

Operating Models

DimensionVertical AI Software (SaaS / Copilot)AI-Native Services (Full Stack / Outcome)
Value PropositionSells tooling to augment internal staffSells the guaranteed end business outcome
Buyer PerceptionEvaluates software UI, latency, securityEvaluates domain credibility, trust, and SLA delivery
Delivery ModelCustomer internal implementationStartup is the delivery team and implementation layer
Primary Failure ModeFeature commoditization by incumbent SaaS"Mirage PMF" (scaling human labor linearly)
Target Gross Margins70% to 85% on day one40% to 60%+ expanding toward software margins via automation

Diagnosing Mirage PMF

Rapid top-line revenue growth and high logo retention can mask an operational failure mode where headcounts scale 1:1 with revenue. Real product-market fit requires proving non-linear operational scaling.

  • Honest COGS accounting: Place all model inference costs, third-party API tokens, and human-in-the-loop review labor directly into Cost of Goods Sold rather than burying them in Operating Expenses.
  • The "HURT" north-star metric: Establish a primary product tracking metric, such as Human Review Time (minutes of human labor required per document or completed task). Margins converge toward software economics as review time approaches zero.
  • ARR per FTE tracking: Monitor revenue per full-time employee specifically across service-relevant delivery staff; if this does not meaningfully outpace legacy industry incumbents, software leverage is missing.

Strategic Entry Wedges

TierTactical ScopeValue DynamicPrimary VulnerabilityDefensible Moat
GoodPoint-solution feature; quick to demoPersonal user convenienceEasily cloned by incumbentsSpeed of execution
BetterEnd-to-end discrete workflow automationHard OpEx reduction & hours savedScope expansion by adjacent toolsDeep workflow data & integrations
BestCore system of record / guaranteed serviceCapacity unlock, margin expansionHigh enterprise frictionHigh switching costs, data gravity, brand

Productization And Delivery

  • Deploy pilot SEAL teams: Staff early customer onboarding with a specialized, highly adaptable team built to navigate operational uncertainty, then hand off successful implementations to steady-state delivery teams.
  • Pair doers directly with builders: Seat domain practitioners (e.g., lawyers, underwriters) directly next to engineers for hourly prompt adjustments and real-time evaluation feedback rather than relying on delayed batch evals.
  • Automate tasks, not whole people: Deconstruct complex roles into discrete, repeatable task units to simplify model evaluation pipelines and ease organizational adoption friction.
  • The 80-client threshold: Resist custom enterprise feature requests until reaching roughly 80 customers, at which point true recurring platform requirements can be distinguished from bespoke edge cases.

Team Building And Domain Authority

  • Borrow instant authority: Hire senior leadership from legacy industry market leaders to provide institutional trust and open specialized hiring pipelines while the AI infrastructure matures.
  • Hire product leadership early: Even when customers never interact with an external UI, product leaders are required to translate messy delivery operations into standardized platform code.
  • Embed domain experts in sales: Include sales engineers or experienced domain specialists in enterprise discovery calls to prevent sales teams from overpromising delivery timelines.

Go-To-Market Strategy

  • Show the internal engine: Ditch slide decks and demo the internal AI tooling in live calls; demonstrating automated work execution in real time cuts enterprise sales cycles substantially.
  • Target standardized mid-market accounts: Start downmarket or in the mid-market where workflow homogeneity enforces productization discipline before accepting complex enterprise custom workflows.
  • Strategic incumbent partnerships: Form distribution alliances with mid-sized legacy service providers, but strictly preserve direct end-customer relationships to protect proprietary data flywheels.

Monetization Playbook

  • Outcome-based pricing: Charge directly per unit of work completed (e.g., per insurance submission processed, per audit binder delivered, per tax return completed).
  • Workload credit systems: Deploy normalized credit models across variable, continuous task streams to decouple pricing from labor hours while capturing task complexity.
  • Value-share benchmarks: Price solutions at 10% to 20% of the customer's total quantifiable cost savings or net-new revenue generated.
  • Transition milestones: If market norms require launching on legacy time-and-materials billing, set clear contractual timelines to transition accounts to outcome pricing as automation scales.

Long-Term Defensibility

  • Contractual data flywheels: Mandate in Master Services Agreements (MSAs) that anonymized engagement data, edge cases, and human-in-the-loop corrections feed proprietary models.
  • System of record lock-in: Build the underlying ERP, ledger, or database of record so that customer historical workflows and data repositories permanently reside on your platform.
  • Deliberate M&A timing: Delay acquiring legacy service providers until the AI platform architecture and automation culture are firmly proven, avoiding the dilution of technology-driven gross margins.

Frequently asked questions

What is Mirage Product-Market Fit in AI startups and how do you diagnose it?

Mirage PMF is the illusion of product-market fit caused by rapid revenue growth and high retention that is powered by hidden human labor rather than true software automation. You can diagnose it if gross margins remain flat or decline (below 50%) as revenue grows, if headcount scales 1:1 with customer onboarding, or if ARR per FTE matches legacy service benchmarks instead of expanding software multiples.

How do you calculate true gross margins and COGS for an AI-native services company?

Unlike traditional SaaS, vertical AI and AI-native service businesses must classify all foundation model inference costs, API tokens, fine-tuning compute, and human-in-the-loop review labor directly inside Cost of Goods Sold (COGS) rather than hiding them in R&D or OpEx. Accurate gross margins typically start at 40-50% during early human-guided workflows and should expand past 70-80% as machine autonomy increases.

What is the HURT metric and how does it measure AI productization?

HURT (Human Review Time) is a north-star product metric that tracks the average minutes of human labor required to review and verify an AI-generated output without degrading accuracy or quality. Pioneer companies like Crosby Legal track HURT on every document processed; as HURT approaches zero, operating margins mathematically converge toward high-margin software economics.

Why should early-stage AI founders start with downmarket or mid-market customers instead of enterprise?

Targeting smaller or mid-market customers enforces productization discipline through workflow homogeneity and lower ACVs, making it impossible to throw unscalable manual labor at delivery. For example, insurance AI startup Harper targeted Main Street businesses (daycares, restaurants) over Fortune 500 accounts, allowing them to standardize document schemas, minimize edge cases, and scale to 5,000+ businesses in 13 months.

How do you structure outcome-based and workload-credit pricing for vertical AI?

Outcome pricing charges directly per discrete deliverable (e.g., $150 per processed insurance claim or $500 per audit binder) rather than per seat or per hour. For ongoing, variable-volume work, use normalized workload credits that map to task complexity, decoupling your top-line revenue from labor hours while capturing 10-20% of the customer's total operational labor savings.

How should AI startups structure Master Services Agreements (MSAs) to build a data moat?

Ensure client MSAs and engagement letters grant your platform explicit contractual rights to retain and train models on anonymized operational data, human corrections, and workflow edge cases. In companies like Harper, every customer interaction and underwriter feedback loop compounds into proprietary matching graphs and domain models that competitors cannot replicate.

How do you run live sales demos for backend AI services without a customer-facing UI?

Instead of relying on conceptual slide decks, live-demo your internal operator workbench and autonomous agent pipelines directly in front of buyers. Enterprise legacy modernization startup Mechanical Orchard cut sales cycles by over 50% by demoing their internal 'Cursor for COBOL' code-refactoring engine live on discovery calls, proving real automation capability to skeptical buyers.

When should an AI startup acquire a legacy service provider versus building organically?

Never acquire a legacy firm early in your lifecycle, as it introduces manual operational bloat and threatens an AI-first engineering culture. M&A should only occur after your core AI platform achieves proven workflow leverage and high gross margins; at that stage, acquisitions serve as an inorganic customer acquisition and distribution accelerant rather than a crutch for missing technology.

What is the 80-Customer Rule for managing custom feature requests in vertical AI?

The 80-customer rule states that founders should avoid rejecting early custom enterprise requests until reaching roughly 80 live clients. As Crosby Legal CEO Ryan Daniels observed, before 80 clients you cannot reliably distinguish genuine platform rules from one-off exceptions; past that threshold, recurring workflow patterns emerge, allowing you to confidently say no to bespoke custom work that does not productize.

How does the Pilot SEAL Team delivery model work in AI implementations?

Staff initial customer onboarding with a specialized 'SEAL team' adept at extreme operational ambiguity and hands-on migration hurdles, rather than rolling general project teams onto pilots. Once implementation stabilizes and data schemas are validated, the SEAL team hands off the account to steady-state delivery operators focused on long-term automation integration.

How should AI startups partner with legacy incumbents without losing distribution leverage?

Partner with mid-sized legacy service providers that have executive buy-in and urgency, rather than slow-moving market giants (e.g., Prosper AI partnering with Firstsource for revenue cycle management). Crucially, startups must maintain direct contractual ownership of the end customer to ensure their proprietary training data flywheels and client relationships are never disintermediated.

Why should vertical AI companies automate tasks instead of whole job roles?

Automating modular tasks rather than entire people makes product adoption palatable to risk-averse enterprise buyers and dramatically simplifies machine evaluation benchmarks. As Strala CEO Timon Gregg noted, focusing on specific task unitslike extraction, reconciliation, or draft generationenables process-oriented engineers to ship reliable, high-precision evals far faster than attempting end-to-end role replacement.

How do you build real-time AI prompt evaluation loops between domain experts and engineers?

Instead of running slow batch evals every few weeks, sit domain practitioners (such as attorneys or underwriters) directly beside software engineers in real-time pairing sessions. At Crosby Legal, lawyers provide granular feedback every few hours, allowing engineers to tweak system prompts and context orchestration in real time, drastically accelerating model performance improvements.

How can vertical AI service providers build sticky recurring revenue like SaaS?

Startups can build recurring SaaS-like stickiness by becoming the foundational system of record or infrastructure layer for the client. For example, AI-native fund administrator Hanover Park built an internal ERP ledger through which all client financial data runs, creating multi-year contract stickiness and prohibitive switching costs.

more than just words|

We're here to help you grow better at every stage of the climb.

let's go to market

Whether you're finding problem-market fit, refining your positioning, shipping product, or scaling go-to-market we're built for every stage of the journey.