Skip to content
WitsCode

AI Chatbot Development Services

Production chatbots on Claude and GPT that answer from your data with human handoff.

4.9from 100+ clients300+ websites shipped, clients in 4 countries

Who this is for

If any of this sounds like you, we should talk.

  • You tried a chatbot widget and it embarrassed you

    A plug-and-play bot gave customers generic answers, or worse, wrong ones. You need a bot grounded in your actual docs, policies, and product data.

  • Support volume is eating your team

    The same 30 questions arrive every day. You want a bot that resolves the repeatable ones and hands the hard ones to a human with full context.

  • You have an AI feature on the roadmap

    Investors, customers, or your own conviction put a conversational AI feature in the plan. You need someone who has shipped LLM products, not someone experimenting on your budget.

  • Your data is the moat, not the model

    Your value is in proprietary content, catalogs, or expertise. You want a chatbot that retrieves from that data securely instead of leaking it or ignoring it.

What changes for you

Outcomes you can point to, not features you can ignore.

  • A chatbot that answers from your knowledge base with sources, and says "I don't know" instead of inventing an answer.
  • Repetitive support conversations resolved automatically, with clean handoff to a human for everything that deserves one.
  • An eval suite of real customer questions with known correct answers, so every release is measured before it ships.
  • Conversation logs and analytics that show what customers actually ask, feeding your content and product roadmap.
  • A bot you own: your infrastructure, your data, your prompts, documented and portable across model providers.

What is included

Scope, organized by phase.

Discovery and data (Phase 1 of 4)

Phase 1

Discovery and data

What we lock down in this phase before moving on.

  • Use-case definition and success metrics
  • Knowledge source audit (docs, policies, tickets, product data)
  • Model selection benchmark (Claude vs GPT on your real queries)
  • Security and data-handling review

How an engagement works

From hello to handoff, step by step.

  1. Scope the job the bot is hired for

    We define what the chatbot should resolve, what it must never do, and how success is measured. No bot gets built without a clear job description.

  2. Pilot on your real data

    Within the first two weeks you get a working pilot answering from your actual content. We benchmark Claude and GPT on your queries and pick with data.

  3. Build integrations and handoff

    We wire the bot into your stack: helpdesk, CRM, live chat, internal systems. Handoff to a human carries the full conversation, so nobody repeats themselves.

  4. Eval, harden, and rehearse

    Every release runs against an eval suite of real questions with known answers. We test for hallucination, prompt injection, and edge cases before customers ever see it.

  5. Soft launch and tune

    The bot goes live to a slice of traffic with human review on every conversation. We tune retrieval and prompts on real usage, then scale to full traffic.

Case study

WitsCode (AI Visibility Checker)

A live LLM-powered tool running in production on witscode.com, generating qualified leads that flow into our CRM automatically. We eat our own cooking before we serve it to clients.

Prospects kept asking the same question: how visible is my brand in ChatGPT, Perplexity, and Google AI answers? Answering it manually meant running queries by hand for every conversation, which did not scale and left no artifact the prospect could act on.

We built the answer as a production AI tool on our own site. The AI Visibility Checker takes a domain, runs live LLM queries, and returns a structured visibility report. The same engineering this service sells: prompt orchestration, API integration, output validation, and lead capture wired straight into our CRM pipeline.

Read the full case study

Why us

What you get with WitsCode that you don't get elsewhere.

  • We ship LLM products, we don't demo them

    Our own site runs a production AI tool with live model calls and real users. We know what breaks after launch because we operate this stack daily.

  • Evals before customers

    Every bot is tested against a suite of real questions with known correct answers. If accuracy drops, the release does not ship. Vibes are not a QA strategy.

  • Model-agnostic, retrieval-first

    We build on Claude and GPT and keep your data layer portable. When a better model ships, you switch in days, not with a rebuild.

WitsCode rebuilt our Shopify store so it finally converts the traffic we were already getting. They understand speed and storytelling in equal measure, and the store has been a real growth lever since launch.
Aravindh NatarajanFounder

Frequently asked

Questions before you reach out.

A chatbot widget wraps a generic model in a chat window and answers from whatever the model already knows, while a production AI chatbot is engineered around your business: it retrieves answers from your documents and systems, follows your rules, escalates to a human when it should, and is tested against evals before launch. The widget takes an afternoon to install and fails on the first real question. The engineered bot takes weeks to build and keeps resolving conversations for years.

RAG (retrieval augmented generation) means the chatbot searches your actual content (help docs, policies, product data, past tickets) before it answers, then generates a response grounded in what it found, which is how you get answers specific to your business instead of confident guesses from a general model. We handle the unglamorous parts that make RAG work in production: chunking your content well, reranking results, keeping the index fresh as your docs change, and citing sources so answers can be trusted and audited.

We build on both Anthropic's Claude and OpenAI's GPT models, and we pick per project based on the task: reasoning quality, instruction following, context length, latency, and cost per conversation all differ, so we benchmark candidate models against your real queries during the pilot and let the eval results decide. We also keep the retrieval and integration layer model-agnostic, so switching providers later is a configuration change rather than a rebuild.

We reduce hallucination with grounding and guardrails: the bot answers only from retrieved sources, cites where each answer came from, declines when retrieval finds nothing relevant, and every release is run against an eval suite of real questions with known correct answers before it reaches your customers. No system makes hallucination impossible, so we also design the failure mode: an uncertain bot hands off to a human instead of bluffing.

A production chatbot typically takes 4 to 8 weeks from kickoff to launch: the first two weeks cover the data audit and a working pilot on your real content, the middle weeks add integrations, handoff, and guardrails, and the final stretch is eval testing and a supervised soft launch. Scope drives the range. A single-purpose support bot lands near the short end, while multi-system integrations with complex escalation logic take longer.

Yes, human handoff is part of every chatbot we ship: the bot detects when a conversation needs a person (frustration, complexity, high-value intent, or an explicit request), transfers with full context to live chat, email, or your ticketing system, and logs the exchange so your team is not starting cold. Handoff is where most off-the-shelf bots fail, and it is the difference between a bot that deflects customers and one that serves them.

Your data stays under your control: we use API access where providers do not train on your inputs by default, keep your knowledge base in your own infrastructure, redact sensitive fields before they reach the model, and scope the bot's read access to only what it needs. We document the full data flow during discovery, so your team (and your customers' legal departments) can see exactly what leaves your systems and where it goes.

After launch we monitor conversations, track answer quality against the eval suite, refresh retrieval as your content changes, and ship improvements on a monthly cadence, either inside a support window included with the build or as an ongoing retainer through our care plans. Chatbots degrade quietly when content drifts, so the maintenance loop matters as much as the initial build.

Ready for a chatbot that knows your business?

Start a project. We will scope the job your bot is hired for and tell you honestly whether AI is the right tool for it.