ProjectsBlogs
AI agent development company — Nexona engineers custom AI agents in code
LangGraph · RAG · Python · Your Infrastructure

AI Agent Development Company

Nexona is an AI agent development company that builds custom AI agents in code — your data, your infrastructure, your repository at the end of it. Not a canvas you rent. We test before we ship.

6-10 wks

To A Production Agent

150-400

Test Cases Before Launch

100%

Code Ownership, Yours

Trusted By

Businesses That Trust Nexona

Custom AI agent development in Python and LangGraph — Nexona engineering work
How We Build

The Demo Takes Four Days. The Agent Takes Eight Weeks.

Custom AI agent development is four jobs, not one: connect the agent to your real data, define what it is allowed to touch, build an evaluation set that proves it works, then run it where you can watch it. The model itself is about a fifth of the effort. Everything wrapped around the model is the rest.

The first demo always lands. Four days, maybe five, and it answers questions in your tone of voice and everyone in the room goes quiet for a second. Enjoy it. That version would fall over in about nine minutes of real traffic.

One build — a support agent for a company selling industrial valves — spent three of its eight weeks on a single problem. Their part numbers looked like VG-4471-B and customers typed them eleven different ways. Spaces, no spaces, lowercase b, the letter O where a zero should be. The agent kept confidently answering about the wrong valve. Not hallucinating exactly, but close enough that it did not matter what you called it.

Fixing that was not a prompt. It was a normalisation layer and 240 test cases written by someone in their sales team who knew every way a customer could get a part number wrong. That is what the eight weeks is.

  • Retrieval agents that answer against your own documents
  • Support and voice agents wired into your live systems
  • Internal copilots for the team, not for customers
  • Multi-agent systems where work is handed between agents
  • Evaluation harnesses so changes can be proven, not guessed
  • Deployed in your cloud account, monitored, and yours to keep

An Agent Is Not A Chatbot

An AI agent takes a goal, picks its own steps, calls tools to carry them out, then checks whether it actually worked. A chatbot answers and stops. Same model underneath — completely different engineering problem, and a completely different set of ways to get hurt.

Chatbot

Answers a question from what it was given. Has no memory of your systems, cannot act, cannot check itself. Useful for FAQs and not much past that. Fails softly — a wrong answer, and the person moves on.

  • —One turn in, one turn out
  • —No access to live data
  • —Cannot take an action
  • —Cheap to build, cheap to be wrong

Agent

Holds a goal across many steps. Reads your live data, calls your systems, decides what to do next, retries when something fails, escalates when it is out of its depth. Fails hard, which is exactly why the guardrails are the build.

  • —Many steps, held state
  • —Reads and writes real systems
  • —Takes actions under explicit limits
  • —Needs evaluation before it goes live

Written In Code. On Purpose.

Build a custom AI agent when the agent is the product. Use a no-code platform when it is the plumbing — genuinely, if your problem is moving a form submission into a CRM, go and use one, you will be live this afternoon. That is a different job and it belongs on our AI automation side. Not sure which of the two you need? That call is part of what a fractional CTO is for.

Canvases stop being the right answer at three points, and they are fairly specific points. When you need to prove the thing works before customers touch it. When per-run pricing meets actual volume. When the logic branches in a way you cannot draw. Also — and nobody puts this in the pitch — you are renting. It runs on their infrastructure, under their pricing, and leaving means building it again.

Evaluated

A written test set of real examples the agent has to pass before launch, rerun on every change. You can see the score. So can we.

Owned

Your repo, your cloud account, from week one. No licence, no per-seat fee, nothing running on our side that you pay to keep alive.

Bounded

An explicit list of what the agent may never do alone. Refunds, contracts, prices, money out — draft-and-approve, always.

Agents We Have Actually Shipped

Retrieval Assistants — AI agent development by Nexona

Retrieval Assistants

Answers drawn from your own documents, with the source attached so somebody can check it. Policy manuals, product catalogues, three years of support tickets. Built as an enterprise RAG chatbot most recently.

Voice Agents — AI agent development by Nexona

Voice Agents

Picks up, understands what the caller wants, checks the real system, books or answers or escalates. Our AI voice agent work — latency matters more than cleverness here, by a lot.

Support Agents — AI agent development by Nexona

Support Agents

Reads the incoming message, pulls the account, drafts a reply, resolves the easy 60% and routes the rest with context attached. See the AI customer support hub.

Browser & Tool Agents — AI agent development by Nexona

Browser & Tool Agents

Agents that operate software the way a person would when no API exists. Slow, occasionally uncanny, extremely useful against legacy portals. The agentic web assistant came out of exactly that.

Multi-Tenant Agent Platforms — AI agent development by Nexona

Multi-Tenant Agent Platforms

One agent system, many customers, strict data isolation between them. Harder than it sounds and unforgiving when it goes wrong — see secure multi-tenant chat.

Internal Copilots — AI agent development by Nexona

Internal Copilots

Not customer-facing. An agent your own team asks — where is this order, what did we quote them last year, draft the follow-up. Often wired straight into your CRM.

See every project

The Decisions That Actually Matter

RAG or fine-tuning?

RAG, nine times out of ten. If the agent does not know your products, prices or policies, that is a retrieval problem — the documents live outside the model, you update them like any other file, and the answer can cite its source. Fine-tuning is for when it knows the facts and still does not sound like you. We have done it twice in three years and undid one.

LangChain or LangGraph?

LangGraph for production, LangChain for the bits around it. LangChain chains steps, which holds up until the agent has to loop, retry, branch or stop and wait for a human. LangGraph makes the run a graph with real state, so you can pause it, resume it, and replay exactly what happened when somebody complains. Sometimes neither — plain Python, no framework, when the job is small. Which is more often than you would guess.

Which model?

Whichever survives your evaluation set, and it changes. We build model-agnostic so swapping is a config change rather than a rewrite, because the frontier moves every few months and you should not be rebuilding each time it does. Cost usually decides it — the cheapest model that passes is the right model.

Where does it run?

Your cloud account. AWS, GCP, Azure, or a box in your own server room if that is genuinely what compliance requires. We deploy into your infrastructure and hand over the keys, which sounds obvious and is not what most of this industry does.

What Drives AI Agent Development Cost

We scope before we quote, so there is no price list here. Three things move the number more than anything else, and none of them is which model you end up on.

Surface Area

How many systems

One job against two systems is the floor. Every extra system the agent has to read from or write into moves the number, and not in a straight line.

Data Condition

How messy it is

Already structured and sitting in a database is cheap. Photographs of printouts are not. Most of the spread in any quote we give comes from this one.

Autonomy

How much it may do alone

An agent that drafts for a human to approve costs less than one allowed to act by itself, because the second needs far more proving before it goes anywhere near a customer.

The one people forget is the running cost. Model usage is billed by volume rather than by seat, so it moves with how hard the agent actually works — which is the opposite of how most software you buy behaves. We put that beside the build price in the proposal. A project that dies in month seven over an API bill nobody mentioned is a project we failed to quote honestly.

FAQ

Agent Questions We Get Asked

An AI agent development company designs, builds and runs software agents that act on their own — reading, deciding, calling your systems, handing off to a person when the job gets unusual. The build itself is mostly unglamorous: connecting the agent to your real data, writing the rules about what it is allowed to touch, testing it against a few hundred real examples before a customer ever speaks to it. The model is maybe a fifth of the work. Everything around the model is the rest.

An AI agent is a program that takes a goal, decides its own steps, uses tools to carry them out, and checks whether it worked. That last part is what separates it from a chatbot. A chatbot answers a question and stops. An agent looks up the order in your database, notices the shipment is late, drafts the apology, applies the credit if it falls under the limit you set, and escalates if it does not. Same model underneath. Very different thing to build, and a very different thing to get wrong.

Four things drive the cost, and the model is not one of them: how many of your systems the agent has to touch, how messy the data is before it is usable, how much of the work is irreversible enough to need approval steps, and whether one agent does the job or several have to coordinate. One agent against a couple of systems sits at the bottom of the range. Multi-agent orchestration sits at the top, and the gap between those two is wide. There is also a running cost — model usage is billed by volume, not by seat — and we put that beside the build price in the proposal, because it is the one that surprises people six months in. We scope before we quote. Anyone who gives you a number before seeing your data is guessing.

Six to ten weeks for a single production agent. The first demo takes about four days and it will be the most impressive the project ever looks — that is the part everyone remembers, and it is roughly 20% of the work. The rest goes on the boring half: edge cases, permissions, what happens when the API times out, what happens when someone asks the support agent for a refund it is not allowed to give. Multi-agent builds run three to five months.

RAG gives the model your information at the moment it answers. Fine-tuning changes how the model behaves. If the problem is that it does not know your products, your prices, your policies, you want RAG — the documents stay outside the model, you update them like any other file, and the answer can cite where it came from. If the problem is that it knows the facts but writes nothing like you, that is fine-tuning. Nine out of ten business projects are RAG. We have fine-tuned twice in three years and reverted one of those.

LangGraph for anything going to production, LangChain for the pieces around it. The difference matters: LangChain chains steps together, which is fine until the agent needs to loop, retry, branch, or wait for a human to approve something. LangGraph models the whole run as a graph with explicit state, so you can see where a run is, pause it, resume it, and replay exactly what happened when a customer complains. We also write plain Python with no framework when the job is small enough, which is more often than framework vendors would like.

Build custom when the agent is the product rather than the plumbing. No-code platforms are genuinely good at connecting apps and firing simple sequences, and if that is your problem you should use one. They get expensive and awkward at three specific points: when you need real evaluation before shipping, when per-run pricing meets high volume, and when the logic needs branching that a canvas cannot express cleanly. You are also renting — the agent lives on their infrastructure under their pricing, and moving it later means rebuilding it.

Yes. About a third of our agent work is embedded — one or two of our engineers inside your team, your repo, your standups. Usually there is already a developer who understands the domain, and what is missing is someone who has shipped agents before and knows which failures are coming. Minimum useful engagement is six weeks. Below that you spend the whole time on context and nothing ships.

It gets things wrong, so the build assumes it. Every agent we ship has three things: a confidence threshold below which it stops and asks a person, a written list of actions it may never take unsupervised, and a log of every run that a non-technical person can actually read. Anything irreversible — refunds, contracts, price changes, sending money — is draft-and-approve by default. We will argue with you if you ask us to remove that.

You own all of it. Code, prompts, evaluation sets, infrastructure config — in your repository and your cloud account from week one, not handed over at the end. We do not keep a copy running that you rent back from us, and there is no per-seat licence. If you want to take it in-house or hand it to another team in eighteen months, nothing about that is our decision to make.

We build the evaluation set before we build the agent. Between 150 and 400 real examples — actual customer messages, actual documents, actual queries pulled from your history — with the correct answer written next to each one. The agent has to hit an agreed pass rate on that set before it goes anywhere near a customer, and it reruns on every change. Without this you are shipping on vibes, and the vibes are always excellent right up until the first week of real traffic.

Access to the data the agent will work from, one person who actually knows the process, and a decision about what the agent is never allowed to do. That is genuinely it. You do not need clean data — we have built against a document set where half the PDFs were photographs of printouts, taken at an angle, in a warehouse. Messy is normal. What stalls a project is the second item: no single person who can say how the process really works, only four people who each know a different third of it.

Tell Us What The Agent Has To Get Right

You do not need a spec. One hour, a description of the job you want handled and the things it must never do on its own, and we will tell you whether an agent is even the right shape for it. Sometimes the honest answer is a database query and four lines of code, and we would rather say that than sell you a model.

Let's Talk
Business

Ready to automate and elevate?

WhatsApp