Agentic AI development company: agents that run in production

Voice agents that answer the phone, analysts that write their own SQL, tools that read contracts and video. What we have shipped, how each one is fenced in, and what it costs to run.

Agentic AI development at ScrumLaunchAgentic AI development at ScrumLaunch

Agents we have built

  • A voice agent that answers the phone

    For home-services businesses: it answers inbound calls, collects the lead's details and transfers to a person when it should. Twilio SIP into LiveKit, a Python agent on AWS ECS, speech-to-text in AWS Transcribe, Claude Haiku 4.5 on Amazon Bedrock in the client's own AWS account, and ElevenLabs voices in English and Spanish. The MVP went from kickoff in mid-March 2026 to production calls on June 9, with 751 unit tests in the voice service and a load-test plan of 100 concurrent inbound and 25 outbound calls.

  • Questions in plain English, answered from the database

    For a professional football club's fan data — communications, ticketing, membership — across 13 tables and more than 11 million communication records. The agent reads the schema and a glossary, writes SQL, validates it, runs it read-only and draws the chart; when a match or a period is ambiguous it has to ask instead of guess. Successful queries become examples for the next ones. Moving it from a DuckDB and S3 pipeline to direct Azure SQL took answers from 40 seconds–9 minutes down to seconds.

  • Two agents over clinical records

    A records platform for clinics that OCRs and indexes charts from EMR systems, so a physician can build patient cohorts and ask about a chart in natural language. One agent, on the Claude Agent SDK, turns a question into safe SQL in up to 15 reasoning steps; another retrieves from the patient's own documents. Neither can step outside the platform's own MCP tools, a network-isolated sandbox and a read-only database role limited to seven tables. It runs live at $0.32 per copilot query.

  • A tool-using analyst over licensed data

    A scout asks "if we sell this player, who similar should we look at in these leagues?" and gets an explainable shortlist. Claude chooses among eight tools — find, profile, search, similar, compare and others — in a loop capped at ten rounds, over about 87,600 player-season rows from 100 leagues refreshed every ten minutes. The model picks and explains; the server computes every number on the cards.

  • Documents and video into structured data

    Purchase orders from four or five customer portals parsed into an ERP upload, with duplicate detection and every calculation kept in deterministic code, delivered as a four-week fixed-price proof of concept for $2,500. And a data-center walkthrough video turned into rack and component records: Whisper on the audio, Claude vision on sampled frames, the two fused into JSON with the evidence and a confidence score for each value, in about 30 seconds for a short video.

  • AI that drafts, and a person who approves

    Inside a HIPAA-aligned product, a trainer generates a four-week program from a member's assessment, then edits and approves it before the member ever sees it; it runs on Bedrock under the client's existing BAA, with no personal identifiers sent to the model. Our own LessonLens does the same for language teachers: evidence-backed observations from a lesson transcript that the teacher accepts, edits or rejects before a student sees a report.

How we build agents that can be trusted with access

  • The model runs in your cloud

    Where we can, the model is called through Amazon Bedrock in the client's own AWS account, through an IAM role, so there are no API keys to leak and the data does not leave an account the client already governs.

  • Deterministic code around the model

    A prompt builder rather than free text, data filtered in code before it reaches the prompt, JSON output validated against a schema, and every number computed by the server rather than the model. The model does what only a model can; everything checkable is checked.

  • Behavior stored as data

    Scenarios, personas and prompts live in the database and are read on every call, so changing how a voice agent handles a call is an edit, not a deployment.

  • Bounded authority

    Only the system's own tools, a network-isolated sandbox, a read-only database role scoped to the tables it needs, a budget of tool calls per request, an iteration limit, and an audit of every attempt to step outside them.

  • Evaluation before model choice

    The same question set runs across several models for quality, cost and latency before one is picked. On the scouting analyst, three Claude models scored 20/20, 20/20 and 17/20 on a 20-question set at $0.015, $0.04 and $0.10 per answer, which makes the choice a trade-off you can see rather than a preference.

  • Sometimes the answer is not an agent

    For invoice entry into a desktop ERP we proposed plain RPA with no LLM in the request path, because the task was deterministic. For a first-version knowledge assistant, full-text search instead of a vector platform. An agent where a script would do is cost and risk without a benefit.

Engagement shapes and what they cost

  • Discovery: 2–3 weeks, $8K–12K

    Which workflows are worth an agent, what data they need, and where a person still has to approve. For one exchange platform this mapped nine use cases, from support and engineering knowledge assistants to reconciliation, AML and surveillance.

  • Proof of concept on your data: 2–4 weeks, $2.5K–16K fixed

    One workflow on real data, with an evaluation set agreed up front. Purchase-order parsing was delivered in four weeks for $2,500; an invoice-automation proof of concept was priced at about three weeks for $12K–13K.

  • Paid pilot: 2–3 months

    The agent in front of real users with limits on. A guided website chatbot was scoped at 290–300 hours with one senior full-stack engineer; three SMS lead agents were priced as a three-month pilot at $5K.

  • An AI feature inside an existing product: about 7–9 weeks

    From scope to first production use. On the training-plan feature, the same dashboard was estimated at 190–280 hours without AI and 380–560 hours with plan generation — the AI roughly doubles the work, which is worth knowing before you budget.

  • Production agent build: 2–6 months, $50K–240K

    Illustrative bands rather than a sum: a knowledge or incident assistant at about $50K–75K over two to three months; an operational workflow $90K–150K; risk, reconciliation or AML $105K–175K; surveillance $160K–240K over four to six months. A voice-agent MVP runs 420–670 hours, about three months with two full-stack engineers.

  • Taking a client-built AI system to production: 8–22 weeks

    For teams that built a working prototype and now need reliability, PHI handling and evaluation discipline: a first phase of 8–12 weeks at $57K–85.5K, and about $114K–156K for 16–22 weeks.

Who builds the agents

Where the team sits

More than half our engineers now sit in Latin America, with most of the rest in Eastern Europe. Agent work runs on an evaluation loop, and someone from the business has to read outputs and say whether they are right; doing that the same day is the difference between a two-week iteration and a two-month one.

How people get here

About ten of our engineers hold Anthropic certifications, and delivery runs on Cursor, Claude Code, Codex and Copilot with human review on every change. Hiring runs through the same funnel as every other stack — roughly one hire per 130 applications on the front end and 157 on the back end.

Frequently Asked Questions

Ready to get started?

Contact us at hello@scrumlaunch.com or fill out the form

get-started-hero

Company Details

Additional Information