Ilyas Ouhnine RAG · Intelligence documentaire

Independent AI engineer — Casablanca · on Paris time

An answer without its source is worth nothing.

The moment the stakes are contractual, a buyer doesn’t want a good answer: they want to verify it in three seconds. I build the systems where that’s possible — provenance carried from OCR all the way to the screen.

Projects from €8,000 · Day rate €500 · available now

Demo — tender file no. 42/2025 · 312 pages · FR/AR scanned
Query

What is the bid bond, and by what date?

Answer

The bid bond is set at 150,000.00 MAD, to be lodged no later than 14 November 2025, 10:00. The required certification is sector E — class 3.

3 values · 3 sources · 0 unsourced

Sources — click a citation

A reconstructed demo built from a typical file, values anonymised — this is not a client document. In production: 274,376 citations across 58 fields, 99.95% of them carrying the verbatim passage.

Measured in production — BidTender, since Sept. 2025

45,000files in production · 56 GB
274,376provenance citations
90.4%accuracy · hand-annotated set
1 mssearch over 113,000 vectors

Why RAG systems break in production

Three reasons, and none of them is the model.

  1. 01

    Real documents are dirty

    Scans, photocopies of scans, tables that exist only as images, structure that changes with every issuer. The pipeline that works on your ten test PDFs does not survive the eleventh.

  2. 02

    Dense retrieval alone doesn’t discriminate

    In a legal corpus, phrasing repeats. Two passages saying opposite things have near-identical embeddings — and what users actually search for, an article number, a date, an amount, is exactly what embeddings handle worst.

  3. 03

    Without citations, the system is unusable

    The moment the stakes are contractual, an unverifiable answer is worth nothing. And traceability isn’t bolted on at the end: it gets lost at chunking. It has to be carried from OCR to the screen.

The pipeline

— what carries provenance from OCR to the screen
01

Ingestion

PDF, Word, Excel, AutoCAD. 40,800 distinct documents.

02

OCR routing

Page by page, bilingual. Only 14% of pages.

03

Chunking

By clause, not by window. 151,661 chunks.

04

Retrieval

pgvector + BM25, HNSW. 113,244 vectors · 1 ms.

05

Citation

The verbatim excerpt, kept. 99.95% of 274,376.

06

Interface

Verifiable in 3 seconds. 4 h → 40 min.

Ways to work together

  1. 01

    Audit and scoping

    A few days to establish why your retrieval is failing, what’s fixable, and what it costs. You leave with a costed plan — whether or not you hand me the build.

  2. 02

    Production build

    The full system: ingestion, retrieval, citations, interface. Shipped with an evaluation set that belongs to you, so you can judge quality without me.

  3. 03

    Ongoing support

    One to two days a month to evolve an existing system, arbitrate technical choices and hand knowledge over to your team.

Got a corpus that fights back?

Describe your documents in three lines — volume, format, what you need out of them. In thirty minutes I’ll tell you whether it’s feasible, how, and what it costs. Reply within 24 hours on weekdays.