AccAnalysisAccAnalysis
Enterprise IT & Cloud

Testing whether prompt-driven business operations actually hold up

A research programme connecting an operational CRM to locally hosted AI models, structured so that competing architectures could be compared on evidence rather than argued about.

At a glance
Status
Delivered
Region
Germany
Engagement model
Joint lab

Status: Delivered. Described from our own delivery records. Client identity, brand and commercial terms are withheld. Figures, where given, cover the period stated and nothing beyond it.

The organisation

Context

A research engagement rather than a production build.

The question was whether business operations on a CRM — the everyday work of finding, updating and acting on customer records — could be driven by prompt rather than by form, and which of the available architectures was actually worth taking further.

The brief

The problem

  • The claims being made about agentic AI on business systems were mostly untested against real operational data.
  • Comparing approaches informally produces an opinion, not a finding: without a fixed task set and a measurement method, whichever architecture is demonstrated last tends to win.
  • Data sensitivity ruled out sending records to hosted models, which narrowed the option set to what could run locally.
  • The cost of being wrong in production is high, and the cost of being wrong in a lab is not.
The work

What we built

An integration between the operational CRM and locally hosted models

Through a tool-server layer, so an agent could read and act on real records under explicit permissions.

Retrieval-augmented storage over the CRM's own data

So answers were grounded in records rather than generated.

A comparison harness

Capable of running the same task set across retrieval-augmented against non-retrieval approaches; several orchestration pipeline variants; and a range of locally hosted models and vector stores.

Method

How we delivered it

  1. 1

    Frame the hypothesis and the success threshold

    Before building — what “good enough to productionise” would mean, agreed in advance.

  2. 2

    Build the integration layer

    Against a masked copy of real data in an isolated environment.

  3. 3

    Establish the task set

    The operations an agent would need to perform, and the correct answer for each.

  4. 4

    Run the comparisons

    Across architectures, models and stores.

  5. 5

    Write up the decision

    What held, what didn't, and what would need to be true to take it further. In writing, including the negative results.

Sequence

How it was phased

PhaseDurationWhat happens
1Framing
1–2 wks

Hypothesis, success threshold, time-box

2Integration layer
3–4 wks

CRM tool server, permissions, masked data environment

3Task set & harness
2–3 wks

Evaluation tasks, expected answers, measurement

4Comparison runs
4–6 wks

Retrieval vs non-retrieval, pipeline variants, model and store variants

5Write-up & decision
1–2 wks

Findings, recommendation, what to productionise

Indicative phasing for work of this shape. Actual duration varies with data quality, access and decision speed.

Hand-over

What the client keeps

  • The prototype and its source
  • The evaluation task set and harness
  • Comparison results for every configuration tested
  • A written decision record covering what was tried and why it was kept or dropped
Stack
Odoo CRMMCP tool serversLocal LLMsLangChainRAG vector storesPython
Keep reading

Related

Related case studies

Tell us what your current system can't do, and we'll tell you what it would take to change that.