Skip to content
Hire engineers

Hire AI developers

The hard part of applied AI is rarely the model. It is grounding it in your data, measuring whether the output is actually good, and designing what happens when it is wrong — which is a product problem before it is a machine learning one.

Grounding, evaluation, and being wrong

A model that answers from general knowledge is a demo. A model answering from your documents, your catalogue or your ticket history is a product, and the distance between them is retrieval, not prompting.

Evaluation is the part most projects skip. Without a way to measure whether output is getting better, every change is a matter of opinion and no one can say whether last week's prompt edit helped. That measurement has to be built alongside the feature.

Then there is being wrong, which will happen. What the interface does at that moment — how confident it looks, what it cites, how easily a person corrects it — determines whether users keep trusting the feature. That is interface design as much as engineering.

What we build, and what we do not claim

We integrate AI into working products: assistance features, retrieval over a customer's own content, and automation of work that was manual. ActivityHub carries an OpenAI integration of that kind; OptionTeller answers user queries through a model served via OpenRouter.

We are not a research lab and do not train foundation models. Being clear about that is more useful than the alternative — if your problem needs original model research, we are the wrong firm and would rather say so at the first conversation.

Suited to
  • Products adding assistance, search or summarisation over their own content

  • Teams who have a working prototype and need it reliable enough to ship

  • Companies automating manual work where being wrong has a manageable cost

Not the right fit
  • Original model research or training foundation models — not what we do

  • Use cases where a wrong answer is unacceptable and there is no human review step

  • Teams wanting AI added because it is expected, without a task it improves

Proof

2 projects built with it

Drawn from the technology each project actually used, so this list cannot include work that did not.

All case studies
ActivityHub case study

EdTech and Activity Management

ActivityHub

A SaaS platform connecting parents with children's activity providers across web, mobile and iPad.

94 min/week
Time saved for activity providers managing activities
863K
Messages sent between users
39K
Sessions hosted on the platform
OptionTeller case study

Trading

OptionTeller

An options analytics platform querying 70+ million daily trades, built on Next.js and Rust.

What drives cost

We do not publish rates, because scope decides them

A figure quoted before scope is understood is not a commitment either side can rely on. These are the things that actually move it.

Data readiness

Whether your content is retrievable, or needs cleaning and structuring first. Usually the largest item.

Evaluation requirements

How rigorously output quality must be measured, which depends on what a wrong answer costs.

Integration depth

A contained feature is far less work than one threaded through an existing product.

Ongoing running cost

Model usage is an operating cost, and architecture decisions move it substantially.

Working together

NDA first, contract on frozen scope

We sign a mutual NDA before the first detailed conversation, not after — you should be able to describe what you are building without qualifying it.

The contract follows requirement freeze rather than preceding it. Once scope is agreed and written down, it is what the agreement is drawn against, which is what makes a fixed commitment meaningful in the first place.

Questions

What buyers usually ask

Something not covered here?

Ask directly — we would rather answer a specific question than publish a general one.

Contact us
  • 01

    Do you train custom models?

    We work with hosted models and retrieval over your data. We do not train foundation models, and if that is what your problem needs we will tell you at the first conversation.
  • 02

    Which models do you work with?

    OpenAI, Claude, Hugging Face and Vertex AI, and routing layers such as OpenRouter where switching model without changing the application matters. Chosen per task rather than by default.
  • 03

    How do you know the output is any good?

    By building an evaluation set alongside the feature. Without one, quality is opinion and no change can be shown to have helped.
  • 04

    What about our data?

    What leaves your systems, where it goes and what is retained are architecture decisions made deliberately and written down.
  • 05

    Can you make our prototype production-ready?

    Often the most useful engagement. The gap between a working demo and a reliable feature is mostly evaluation, error handling and cost control.
Next step

Tell us what you want it to do

Describe the task and what a wrong answer would cost. We will come back with an approach, and say plainly if AI is not the right tool.

  • NDA before we talk
  • Reply within one business day
  • No obligation