Retrieval before generation
Grounded in your documents, your tickets, your schema — with citations the user can open. An answer with no source is a guess wearing a suit.
AI
An assistant in production needs three things a demo does not: a bounded set of actions, a record of what it did, and a way to measure whether it is getting better or worse.
Softrear builds LLM assistants and agents for production use, with retrieval grounded in the client’s own data, explicit tool permissions, human approval on consequential actions, and an evaluation suite that runs in CI.
Grounded in your documents, your tickets, your schema — with citations the user can open. An answer with no source is a guess wearing a suit.
Every action the agent may take is declared, permissioned and logged. Consequential actions get a human in the loop by default, not as an upgrade.
A graded test set that runs on every change, so a prompt edit that quietly degrades accuracy fails a build instead of a customer conversation.
When confidence is low the system says so and routes to a human. Confident wrongness is the failure mode that destroys trust in these systems.
Token spend and response time tracked per feature, with caching and smaller models used where they are indistinguishable to the user.
Guardrails
Almost every stalled AI project we are called into has the same cause: the data exists, but nobody can say where it came from, when it was last correct, or who is allowed to see it.
The cheapest way to find out whether an AI project is worth doing is to spend two weeks finding out, with people who will tell you if the answer is no.
A senior engineer reads every one of these. If we’re not the right fit we’ll say so, and point you at someone who is.
No spam, no drip sequence. A senior engineer replies within one business day.