Answers from a manufacturer's catalogues
Manufacturer · pilot on staging
Hundreds of PDF pages turned into a chatbot that answers with the document and page it took the answer from. Built as a pilot: the numbers below decide whether it goes further.
[Services / 02.3]
◆ machine · generated · 24 tokens · 1620 ms
Before choosing a model, I write down the questions it must answer, the actions it may take and the point where a person steps in. Then I build the smallest pilot that can fail honestly.
The demo chatbot invented a discount code.
fix → Answers are grounded on the client's own content, cite their sources, and say “I don't know” when nothing fits.
Nobody can tell if it's getting better or worse.
fix → An eval set of real customer questions from day one, re-run on every change, with scores in a written report.
The API bill tripled in a month.
fix → A monthly cost cap, cost per conversation in the eval report, and an alert before the budget runs out.
Legal blocked it a week before launch.
fix → Data flows mapped at kick-off: what goes where, what never leaves the client's systems, where the models run. Written down, ready for the DPO.
Assistants that answer from the client's catalogue, policies and docs, cite the page they used and hand over to a human when they should.
Search that understands “waterproof jacket for the Pyrenees in March”: hybrid keyword and vector search across products, articles and knowledge bases.
Technical sheets, price lists and manuals turned into answers that point to the document and page they came from.
Tagging, translation, product descriptions, alt text and internal search in WordPress and WooCommerce, inside the editor the team already uses.
Test sets, refusal rules, prompt-injection checks, cost and quality dashboards. Also as an audit of an AI feature someone else built.
Every question is matched against the client's own content first. Only passages above a similarity threshold reach the model, the answer cites them, and when nothing clears the bar the assistant says so and calls a person.
The third question is the one that matters. A useful assistant isn't the one that always answers. It's the one that knows when the sources aren't enough.
01 · Retrievethreshold 0.75
02 · Answermodel: any
Yes. Orders to the Canary Islands arrive in 5–7 working days, and customs are already included in the price [01]. Free delivery over €60 applies to mainland Spain only [02].
sources ≥ 0.75: 2 · cited: 2 · mode: answer
Yes, within 14 days, as long as it's unworn and still has its tags [03]. The refund goes back to your card within 5 working days [04].
sources ≥ 0.75: 2 · cited: 2 · mode: answer
I don't have a reliable source for that, and I'd rather not guess about safety gear. Want me to pass you to someone in the shop?
sources ≥ 0.75: 0 · cited: 0 · mode: handoff to a human
Which questions or tasks, which content answers them, where it lives and what must never leave the client's systems.
A working version on the real content, on a staging link, plus an eval set built from real customer questions.
Accuracy, citations, refusals, latency and cost per conversation, in writing. Go, adjust or stop, with numbers.
Into the site, app or CMS: accessible UI, guardrails, analytics and a human handoff that actually reaches someone.
Evals on every change, cost alerts, a monthly review of failed answers. Documentation and no lock-in to one model.
Manufacturer · pilot on staging
Hundreds of PDF pages turned into a chatbot that answers with the document and page it took the answer from. Built as a pilot: the numbers below decide whether it goes further.
Every pilot ends with this report
Retrieval-augmented generation: before the model answers, the system looks up the most relevant passages in the client's own content and gives them to the model with the question. The answer is built from those passages and can cite them, instead of coming from the model's memory.
It's the main risk, and it's managed, not wished away. Answers are grounded on retrieved sources, the assistant is told to decline when nothing relevant is found, and an eval set checks this on every change. Where a wrong answer is costly, it hands over to a human.
Whichever fits the task, the language and the budget: usually Claude or OpenAI models, sometimes open models running on premise. The integration is built so the model can be swapped without rebuilding the product.
We map it at kick-off. Where the models run is decided with you, project by project, based on the data involved. Personal data stays out of prompts unless it's needed and documented.
There's no price list for the build: every project is scoped. Running costs depend on traffic and model, so the pilot measures the cost per conversation on real questions, and production gets a budget with alerts.
Yes. Search, an assistant or editorial tools can be added to a live WordPress or WooCommerce site as a plugin or a separate service, without rebuilding the theme.
When the project needs it, yes. I build AI features that agencies offer their own clients, under the agency's name and contract. Companies with an in-house product or marketing team can also work with me directly.
The ones your client's customers actually ask. They become the first eval set, and the fastest way to see if AI is worth it.