Technology
How Gilbert reaches 98.2% accuracy in French
Training a speech recognition model for French is not a matter of translating an English one. A look at our technical approach.

98.2% accuracy in French. We did not invent that number: it is measured on an internal benchmark of 500 hours of real professional conversations across ten sectors. Here is how we got there.
Why French is a particular challenge
Most speech recognition models are trained mainly on English, then adapted to other languages. That approach causes several problems in French:
- Liaison and elision: "les entreprises" is pronounced as one block, not as two separate words
- Homophones: "compte", "conte" and "comte" sound identical
- French technical vocabulary: legal, financial and medical terms have specific pronunciations that English-trained models have never met
- Regional accents: the same word can sound very different in Marseille, Lille or Paris
Our approach: native training, not translation
We did not take an English model and fine-tune it for French. We built our acoustic and language models natively for French, in several layers:
1. A native French acoustic model
Trained on thousands of hours of French audio, and specifically on professional conversations rather than podcasts or audiobooks. The difference matters: the rhythm, the interruptions and the background noise of a video call look nothing like a studio recording.
2. Sector language models
For each professional vertical, from legal to finance to technology to healthcare, we assembled corpora of specialised vocabulary. When a lawyer says "clause résolutoire", our model knows that is more likely than "close résolutoire". This sector context is what takes accuracy from 92% to 98%.
3. Advanced diarisation
Knowing who is speaking matters as much as understanding what is said. Our diarisation handles two to fifteen simultaneous speakers, even when they talk over each other, which happens on average every 47 seconds in a French professional conversation.
The benchmark: 500 hours, ten sectors
Our benchmark covers:
- 50 hours of leadership meetings
- 60 hours of legal conversations
- 45 hours of investment committees
- 55 hours of consulting workshops
- 40 hours of sales calls
- 50 hours of engineering rituals, from sprint reviews to post-mortems
- 40 hours of HR interviews
- 35 hours of clinical meetings
- 35 hours of property valuations
- 40 hours of compliance audits
Across the whole set, the average word error rate is 1.8%, which is 98.2% accuracy. In sectors such as finance and legal, where vocabulary is more structured, we reach 98.7%.
What comes next
Our target is 99% by the end of 2026, mainly by improving how we handle regional accents and bilingual conversations, French and English mixed, which are increasingly common in business.
Accuracy is not a marketing number. It is the foundation of trust. When one technical word is mistranscribed, the whole chain collapses with it: the summary, the decisions, the actions. That is why we invest so heavily here.



