Integrations, Data & AI

An AI agent runtime with a leash
tools, limits and a human approval step built in

An AI agent that can only talk is a chatbot. An agent that can update a CRM, check inventory, or issue a refund needs more. It needs a runtime that limits what it can touch and asks a human before anything irreversible. We build that runtime, not just the conversation on top of it.

from$2,000
Timeline2 to 6 weeks
What is includedTool definitions scoped to exactly what the agent needs, nothing broaderA human approval step for any action that is costly, irreversible, or outside normal boundsFull logging of every tool call, its input and its resultA kill switch that stops the agent instantly across every channelRate limits and spend caps per tool, independent of the LLM gateway's own budget
2-6 weekstypical time from kickoff to an agent running real tools in production
everyirreversible action gated through a human approval step, not just logged after the fact
instantkill switch across every channel the agent runs on

What this layer actually does

An AI agent runtime is what lets a language model do something beyond generating text. Call a function that checks inventory. Update a CRM record. Issue a refund. Send a message.

It is also the layer that decides what the model is allowed to do. What needs a human to say yes first. What gets logged so a decision can be reviewed later.

The conversation people see is the smallest part of this. The runtime around it is what makes an agent safe to run in production: permissions, approvals, logging, a kill switch.

Where the risk actually starts

You need this once an AI agent needs to take action, not just answer questions. Checking real inventory. Updating a real order. Issuing a real refund. That is exactly where a confident but wrong model response stops being an annoying chat answer and starts being a real-world mistake.

It is also essential once several channels, a website chat, WhatsApp, Telegram, all need the same agent with the same limits. Otherwise each channel’s integration ends up reinventing its own, inevitably inconsistent, safety rules.

You do not need a full runtime for an agent that only answers questions from a knowledge base, with no ability to change anything. That is a RAG pipeline, a simpler and cheaper build. The runtime earns its cost specifically once “the agent can do something” becomes true.

How we build it

Every tool the agent can call is defined narrowly and explicitly. An agent that can “look up an order” gets exactly that function, not broad database access it could misuse in a way nobody anticipated.

We define approval rules with you before launch, not as a generic default. A refund under a set amount might go through automatically. A refund over it waits for a human. A price change always waits. The specific lines depend on your actual risk tolerance, not ours.

Every tool call, its input, and its result get logged, so a reviewed decision later has the full trail, not just the final chat transcript. A kill switch stops the agent across every channel at once, built in from day one rather than added after an incident makes the need for one obvious.

We built this exact pattern into a seven-channel sales agent, now running hundreds of unit tests against its tool-calling logic. We also built it into the agent team running day-to-day operations for ProBay, the marketplace we are launching.

What to watch

The biggest risk in an agent runtime is scope creep. A tool added quickly for one use case can turn out to allow something nobody intended, once the agent finds a creative way to use it. We review tool scope specifically for this before each addition, not just whether the tool works.

The second risk is approval fatigue. If everything requires a human, the agent adds overhead instead of saving time. Approval rules need real calibration against actual risk, not a blanket “ask a human for everything” default that defeats the purpose.

Cost of ownership is mostly maintaining the tool definitions and approval rules as your business changes. A new refund policy or a new product category means revisiting the rule set. That is a normal part of running an agent, not a sign something was built wrong.

We also test the agent against adversarial inputs before launch, a user deliberately trying to get it to do something it should not. That test tends to surface a permission gap that testing with cooperative users alone never will.

What it costs

Scope Price Timeline
2-3 tools, basic approval flow from $2,000 2 to 3 weeks
Multiple tools, multi-channel, full escalation from $5,000 4 to 6 weeks

Where this connects

Built as part of AI agents and custom development. Runs on top of an LLM gateway with cost control and benefits from a model evaluation and test harness before launch. See it in production in a seven-channel AI sales agent, and in the agent team we are building for ProBay. Tell us what you want an agent to actually do: get in touch.

FAQ

How much does an AI agent runtime cost?

A runtime with two or three tools and a basic approval flow starts at $2,000. A fuller build with many tools, multiple channels and a detailed escalation workflow runs $4,000 to $10,000. That is close to what we quote for a full AI agent build, since the runtime is most of that work.

How long does it take?

2 to 6 weeks, depending on how many tools the agent needs and how much judgment the approval rules require. A narrow agent with one or two well-defined tools ships faster than one touching several systems with different risk levels.

What counts as needing human approval?

Anything costly, irreversible, or outside a normal pattern. A refund over a threshold. A price change. Deleting a record. An action the agent has not seen before. We define the exact rule set with you before launch. It is a business decision, not a default we impose.

What stops the agent from doing something it should not?

Scoped tool permissions first. The agent cannot call a tool it was never given access to, no matter what it decides to try. Approval rules second, for things it is allowed to touch but should not do without a human checking. A kill switch third, for everything else.

Who owns the runtime and the logs?

You do. It runs on your infrastructure. Every tool call and its result is logged on your systems. The agent's permissions are defined in your own configuration, not inside a vendor's black box.

Start here

Tell us the problem.
We bring the system.

A 30-minute call, then a written plan with numbers within 48 hours. No obligation. If we are not the right fit, we will say so and point you to someone who is.

LIKE WHAT YOU SEE?

This site is our work.
Want one like it?

Ten languages, no page builder, launched in 2026 by a team working since 2015. We can build the same quality into your site.

  • 10 languages
  • Since 2015
Get a site like this →