Case studies

Systems I've built that run every day

Ten production systems from Hariprasad Sivakumar, a GTM engineer in Bengaluru. Most were built AI-native with Claude Code for a GTM agency running outbound for 10+ B2B clients, and they run without anyone babysitting them. Every number below comes from the build logs, not a pitch deck.

AI agent · Slack · Notion

An AI GTM analyst that lives in Slack

A Slack assistant the whole GTM team talks to in plain English, answering from live data instead of memory.

The problem

An agency running outbound for 10+ clients had its answers scattered across sequencer workspaces, Notion, Slack threads and half a dozen provider dashboards. Every "how is this client doing?" or "is this campaign ready to launch?" meant someone opening five tabs, and weekly reporting ate hours.

What I built

  • One tool-calling agent routes every message. Code renders the numbers as Slack cards and the model adds a line or two at most, so it can't invent a metric. Unknown is reported as unknown, never as zero.
  • Weekly client reports comparing this week with last (reply rate, positive reply rate, meetings per 1,000 sent, bounce rate) with a PDF you can re-layout from the thread: "remove the chart, add a summary".
  • Positive-reply feed: replies flow from the sequencer through Clay into Slack, the agent re-checks the sentiment with an LLM, finds the prospect's phone through a provider waterfall and tags the growth team. Follow-ups thread under the original alert.
  • Deliverables tracker: doc links, files, client status updates and purchases posted in Slack are logged into Notion automatically, with a quiet check mark and one-tap undo.
  • Infra Q&A: blacklist status per domain, renewals, stealth-redirect status, launch preflight ("is X ready?") and money leaks.
  • Guarded writes: every action shows a preview and waits for a typed "yes", enforced again by the backend.
  • Report hub: a daily snapshot of every sequencer workspace, a dashboard, and PDFs rendered with headless Chromium.

Results

  • A deep QA pass found and fixed 82 wrong answers (outbound, infra and preflight) before the team relied on it.
  • The AI re-check caught 5 of 25 replies tagged "positive" that were actually misreads.
  • Moving reply alerts from a 5-minute poll to a zero-poll listener removed ~29,000 API calls a day.
  • 16 live agent regression cases and CI house rules (no secrets, consistent style) run on every change.
Cold email infrastructure · Control plane

A cold email command center for a 10+ client operation

One place to run campaigns, inboxes, domains and deliverability across every client and every vendor.

The problem

Outbound at agency scale means hundreds of domains and thousands of inboxes spread across registrars, inbox providers and the sequencer, each with its own dashboard. Problems like a broken DNS record, a disconnected inbox or a campaign starved of sending capacity only showed up after sends failed.

What I built

  • A multi-client control plane with a health score for every inbox and every domain tracked from registrar to campaign.
  • 10 provider integrations: Smartlead, Zapmail, Hypertide, InboxKit, Spaceship, Dynadot, GlockApps, Slack, OpenAI and a stealth-redirect layer.
  • A deliverability doctor, positive-reply attribution, lead recycling, and client onboarding and offboarding flows.
  • Strategy, Copy and Spintax studios that enforce copy that reads human in every combination.
  • A natural-language command palette and an MCP server sharing one tool registry. Safe mode is on by default and purchases are blocked in code.

Results

  • Grew from 71 campaigns, 922 inboxes and 49 domains at the first build to 5,463 inboxes and 404 domains in a single workspace.
  • Campaign detail went from 12s+ to ~1.1s, the domains view from 80s to 3s, and one bulk call replaced 71 analytics calls.
  • Placement monitoring narrowed 3,045 candidate inboxes to 20 test slots that still cover every client and both email providers.
AI · Model Context Protocol

A 100+ tool MCP server: run outbound from Claude

The entire command center, operable from Claude Code, Codex or claude.ai in plain English.

The problem

Every new question or fix meant building another screen. The team already lived in AI assistants, so the operation should be reachable from there, safely.

What I built

  • An MCP server exposing the platform's tools over the same registry as the in-app command palette.
  • Reads run instantly. Writes return a preview, offer up to three ranked options with a recommendation, and wait for an explicit pick.
  • Hard guardrails: it never purchases and never starts a campaign on its own.
  • A hosted Streamable HTTP endpoint with OAuth 2.1, PKCE and dynamic client registration, plus keyless list-building tools.

Results

  • 79 tools at launch in July 2026, 100+ today: full platform parity except payments.
  • Any Claude or Codex session can audit inboxes, build a campaign draft or answer an infra question without opening a dashboard.
Email infra · Orchestration

Domain-to-inbox provisioning orchestrator

From a client's website to live, connected sending inboxes as one tracked workflow.

The problem

Standing up sending infrastructure for a new client was a checklist across a registrar, DNS, two or three inbox providers and the sequencer: days of manual steps, each easy to get subtly wrong.

What I built

  • A state machine: brand-safe AI domain ideas, purchase hand-off (always a human checkout), ownership verification, Google and Microsoft split, nameservers, provider connect, forwarding, persona inboxes, live.
  • Retries, idempotency keys, an audit log and manual override at every step.
  • AI domain generation returns about 15 available, on-brand ideas in roughly 5 seconds.

Results

  • First client run: 20 domains became 156 inboxes (48 Google, 108 Microsoft), all connected to the sequencer.
  • A 500K-emails-a-month build: 180 domains pointed and verified and 2,220 new inboxes, about 504,000 emails a month of modeled capacity.
Email infra · Self-healing

Deliverability Guard + Campaign Autopilot

A sender fleet that checks itself every morning and rebalances itself every day.

The problem

Inbox fleets decay quietly. DNS drifts, inboxes disconnect, warmed inboxes sit idle, and campaigns keep sending from broken senders until reply rates fall off a cliff.

What I built

  • Deliverability Guard: a daily pre-send job that scans every domain and inbox, auto-fixes DNS, pulls broken inboxes out of live campaigns and posts a short Slack digest grouped by domain.
  • Campaign Autopilot: locks each campaign's sender persona, graduates newly warmed inboxes in using provider-aware rules, prunes and replaces broken senders, and sets leads per day from real sending capacity.
  • Drain-safe removal, so a sender with follow-ups still in flight is held, never orphaned. Both run report-only until armed.

Results

  • Piloted on a 4,191-inbox client aiming for 500K emails a month.
  • Found 370 warmed inboxes ready to add and 591 safe un-shares, while deferring 69 inboxes with mail in flight.
  • The first live run pulled 102 broken senders and held 51 that still had follow-ups queued.
Data · Lists · Quality

Lead-list factory with a 5-step QA gate

Signal-based lists that are verified before they're delivered, every time.

The problem

Lists fail quietly: outdated titles, catch-all emails, people who changed jobs, contacts in the wrong country. Every bad row is a bounce that costs domain reputation.

What I built

  • Sourcing from real signals: job posts, tech stack, directories and company databases.
  • An enrichment waterfall across data providers, then suppression against blocklists, past responders and active campaigns.
  • A mandatory QA gate: title recheck against persona definitions, field normalization, two-stage email verification (MillionVerifier then BounceBan), email-to-company affiliation match, and person geography.

Results

  • 100,000+ contacts sourced across client builds in 2026.
  • A sales-tool campaign built from 2,017 companies with a hiring signal: 1,816 contacts at 1,162 companies, 100% verified emails, 1,157 with a phone.
  • A six-vertical CTV advertising build of 44,867 leads with zero duplicates.
  • On its first run the gate rejected 658 of 2,589 rows that would otherwise have shipped.
AI · Multi-agent research

Prospect research agents for private credit

Four trained Claude subagents that qualify accounts the way an analyst would, with sources.

The problem

For a fintech client selling to lenders, the real qualifier was whether a company runs an asset-backed debt facility. No database has that as a filter, so firmographic lists were mostly noise.

What I built

  • Four specialized agents, one per segment (private credit funds, fintech lenders, specialty lenders, banks and neobanks), weighted to the client's priorities.
  • Every company must prove the qualifier with a source URL, is deduped against an 8,500+ row suppression set, and is capped at a daily ceiling, not a quota.
  • A self-learning ledger, so each run gets better at finding what worked.

Results

  • Bank segment: 60 candidates re-verified by QA agents into 50 sponsor banks and 43 neobanks, 743 contacts.
  • A recycle round re-activated 4,108 leads after syncing the do-not-contact list.
Signals · Technographics

Free tech-stack signals from DNS

Proof of what a company actually uses, read from its own DNS records at no data cost.

The problem

Technographic data is expensive, often stale, and can't answer newer questions like "does this company already pay for an AI model?" or "do they run Teams or Slack?".

What I built

  • A scanner that reads SaaS verification TXT records, MX and SPF to fingerprint a company's stack.
  • A platform gate that classifies Google Workspace vs Microsoft 365 before a Teams- or Slack-specific campaign.

Results

  • 1,909 domains scanned in 2 minutes 34 seconds.
  • Across 1,910 US software companies with 50 to 250 staff: 46% verified with Anthropic, 24% with OpenAI, 14% with Cursor, and 53% with at least one AI vendor.
  • An MX gate across 30,225 domains measured a 3x spread in Microsoft share by industry (42.9% vs 14.9%), which reshaped targeting.
Signals · Hiring intent

OpJobs: hiring-signal engine

A company hiring for GTM roles is a buying signal. OpJobs catches it the week it happens.

The problem

Job boards are noisy and full of duplicates. The same role shows up across boards and cities, and keyword searches return mostly irrelevant postings.

What I built

  • Collection from public job board pages only, deduped across boards and cities.
  • Rule-based enrichment plus an optional Claude "why now" angle for each posting.
  • Only net-new jobs are pushed into Clay, with daily monitors keeping it running.

Results

  • Relevance tuning took results from 1 relevant posting in 40 to 40 out of 40.
Product · Clay alternative

hflow sheets

A spreadsheet-native GTM workspace: the tool I wished existed while running enrichment at scale.

What it does

  • Waterfall enrichment across 10 data provider integrations, with verified email finding.
  • 30 column types, including AI research agents that fill a column from the web.
  • Company discovery from maps and lookalike search, all inside a familiar sheets interface.
Also shipped

More builds

Ad-intelligence engine

Active Meta, Google and LinkedIn ads for any domain, batch-run across 12,500 domains, so outbound can open with what a prospect is running.

Ops monitor hub

Live credit balances for 7 data vendors, a daily blacklist and DNS diff to Slack, and a cost ledger for rented inboxes with AI invoice extraction.

Landing pages for outreach domains

A real page on every sending domain, served with on-demand TLS and provider plugins, so prospects who click through find something credible.

Multilingual AI voice assistant

Hands-free, ChatGPT-style voice mode in a production app: Whisper in, LLM reasoning, natural speech out, answering in the user's language.

Realtime finance platform

Automated pixel-perfect PDF invoicing, a realtime sync hub and full offline capability. Systems only count if they keep running.

Encrypted knowledge vault

A desktop app for client knowledge, built with Tauri and Rust, with Argon2id key derivation and XChaCha20 encryption at rest.

Want one of these in your stack?

Big budget or lean, if the problem is interesting, let's talk.

Book a strategy call