Fixed scope
Internal knowledge assistant
Hybrid retrieval across scattered internal documentation, with citation enforcement and an eval harness that gates every prompt change in CI.
Services / AI, ML & MLOps
Most AI pilots die in month two. We build the parts that decide whether yours survives.
The problem
The demo works because someone asked it three friendly questions. Production breaks on the fourth kind.
So we build an evaluation set from real user questions first, then measure every retrieval, prompt and model change against it. You get a number, not an opinion.
In the work
Dense and sparse retrieval fused, re-ranked, with a floor below which the system declines to answer. The eval gate blocks the deploy when faithfulness drops.
Capabilities
9 capability groups
Typical engagements
Indicative scope and duration
Fixed scope
Hybrid retrieval across scattered internal documentation, with citation enforcement and an eval harness that gates every prompt change in CI.
Fixed scope
A multi-step agent that reads a request, gathers what it needs from your systems, drafts the action, and stops for human approval where the cost of error is real.
Advisory
Independent review of a pilot that is not converting: retrieval quality, prompt architecture, latency budget, unit economics, and whether the use case is winnable at all.
What you get
A versioned eval set with pass thresholds wired into CI, so quality regressions fail a build rather than a customer conversation.
Documented chunking, embedding, hybrid search and re-ranking decisions with the trade-offs and the benchmark numbers behind each.
Cost per query and p95 latency at your projected volume, with the levers that move them ranked by effort.
What to do when quality drops, retrieval goes stale, or a provider deprecates a model — written for your on-call engineer.
Questions
Usually not first. In most engagements retrieval quality, prompt architecture and re-ranking move accuracy far more than fine-tuning, at a fraction of the cost and with none of the retraining burden. We fine-tune when there is a measured ceiling that retrieval cannot lift — typically format adherence, domain vocabulary, or latency-driven use of a smaller model. We will tell you which case you are in before you spend on it.
Yes. We deploy open-weight models on your own infrastructure with vLLM or Triton when data residency, contractual restrictions or cost make hosted APIs unworkable. The trade-off is real — you take on GPU capacity and evaluation burden — and we will quantify it before you commit rather than after.
Three layers, none of which is a prompt asking it politely. Retrieval returns scored passages and the system refuses when scores fall below a floor. Answers must cite retrieved spans, and uncited claims are stripped. The eval harness includes questions with no correct answer, and refusing them correctly is a passing result.
Also from Windsor Harlow
Governor-safe Apex, tested triggers, automation a new admin can read.…
→ →Cost and operability designed in as requirements, not cleaned up after the invoice arrives.…
→ →Distributed patterns that survive growth, without the premature architecture that sinks a product first.…
→Tell us the system, the constraint, and what happens if it is not solved. A senior engineer replies within one business day.