The problem
Teams don't keep their knowledge in one place. It's in PDFs, in a competitor's website, in a two-hour recorded workshop nobody has watched twice, in a founder's podcast appearance. Making that usable means ingesting formats that behave nothing like each other, and then answering questions against them without inventing things.
There was a commercial problem underneath the technical one. Retrieval-augmented systems consume tokens unpredictably. A platform that lets users point at arbitrary content and start asking questions can develop a cost structure that quietly destroys its own margin.
Anyone can demo RAG. The engineering is in making it cheap, grounded and predictable when thousands of people use it at once.
What we built
We engineered the backend and retrieval architecture behind a node-based canvas, where users connect content sources to chat and assistant nodes and build workflows visually.
- Multimodal ingestion. Web and social scraping, document parsing, and audio/video transcription — each source normalised into the same downstream representation.
- Retrieval architecture. Vector storage and retrieval tuned so answers are grounded in the source material and traceable back to it, rather than plausibly worded guesses.
- A unified token-metering layer. Usage measured and attributed per workspace, so cost is visible, predictable and controllable rather than discovered at the end of the month.
- Workspace isolation. Each customer's content, indexes and usage kept properly separated at the infrastructure level.
How the agents were used
This engagement is where the AIOBC Knowledge Agent came from. The ingestion, transcription, retrieval and grounding pipeline we built here became the reusable component we now deploy — configured to each client's sources — in a fraction of the original build time.
We aren't faster because we cut corners. We're faster because someone already paid for the hard version, and we kept the engineering.
The outcome
The platform shipped with retrieval that stayed grounded and costs that stayed forecastable — the two things that most often break AI products between demo and production. The backend infrastructure and retrieval architecture formed part of the technical foundation of a platform that went on to raise $75M in venture funding.
What we’d tell you before you build one
Most teams asking for “a chatbot over our documents” need about a fifth of what they think they need, and they need one thing they haven't thought about at all: a way to tell whether the answers are right. Evaluation isn't a nice-to-have you add later. It's how you find out, before your customers do, that the system has started confidently making things up.