How to Build an AI Knowledge Base for a UAE Business
How to build an AI knowledge base for a UAE business: bilingual Arabic and English content, chunking, hybrid retrieval, citations, access control and residency.
Quick answer
An AI knowledge base is your business's approved documents and data, prepared so an AI assistant can find the right passage and answer from it with a citation. For a UAE business, build it in six steps: choose a use case, collect and clean bilingual content, parse and chunk it, index it for hybrid search, generate cited answers, and escalate when sources are missing.
The UAE-specific work is in the detail: Arabic and English versions of the same policy, scanned Arabic documents, Arabic spelling variants, data protection rules that differ between the mainland, the DIFC and ADGM, and health data that must stay in the country. This guide focuses on those points and links to our RAG guide and the rest of the retrieval cluster for technical depth. For a shorter, general introduction, read AI knowledge base: building an assistant on company documents.
Key takeaways
- Start with one use case and one audience; the content and permissions follow from that.
- Leave out secrets, unapproved prices, stale drafts and personal data without a lawful basis.
- Use provider chunking defaults as starting points (800/400 tokens at OpenAI, 512 tokens with 25% overlap in Azure guidance), then test.
- Combine keyword and vector search; keyword search catches product codes, names and dates that vectors miss.
- Normalise Arabic (alef forms, ta marbuta, alef maksura, diacritics, tatweel) so spelling variants match.
- Decide which language version is authoritative, and keep parallel versions in sync.
- Enforce document-level permissions in retrieval, not in the prompt.
- Evaluate with your own bilingual test set; do not rely on vendor benchmarks.
What an AI knowledge base is, in exact terms
Definition: an AI knowledge base is a governed collection of approved content (documents, pages, records and structured data), converted into a searchable index so that an AI system can retrieve the most relevant passages for a question and generate an answer grounded in them. The pattern behind it is retrieval-augmented generation (RAG).
It differs from a traditional help centre, which customers search themselves, and from a chatbot with scripted answers. It is also different from training a model: the model does not memorise your documents; it reads the retrieved passages at answer time. That is why it can stay current and cite sources. For when training does make sense, see RAG vs fine-tuning.
| Term | Definition |
|---|---|
| Ingestion | Collecting source content and converting it into clean text with structure and metadata |
| Parsing / OCR | Extracting text, headings and tables from files; OCR (optical character recognition) reads text from scanned images |
| Chunk | A passage of a document, sized for retrieval, stored with metadata such as title, language, owner and date |
| Embedding | A list of numbers representing a chunk's meaning, so similar meanings sit close together |
| Vector search | Finding chunks whose embeddings are closest to the question's embedding |
| Keyword (BM25) search | Ranking chunks by matching words, weighted by how rare and frequent they are |
| Hybrid search | Running keyword and vector search together and merging the results |
| Reranking | A second model re-scoring the top candidates for relevance to the question |
| Grounded generation | A model writing the answer only from retrieved passages |
| Citation | A pointer from a claim in the answer to the passage that supports it |
| Security trimming | Removing results the user is not allowed to see before they reach the model |
Why UAE businesses need one
UAE context. Generative AI use is very high: Microsoft's AI Economy Institute estimates 70.1% of the UAE working-age population used a generative AI product in Q1 2026 (Microsoft). An AWS and UAE AI Office study reports 72% of UAE businesses have adopted AI (Zawya). When staff and customers ask AI tools about your business and the tools have no access to your approved content, they guess, or staff paste company documents into consumer tools.
Three practical reasons. First, consistent answers: support, sales and operations staff answer from the same approved source, in English and Arabic. Second, faster onboarding in a workforce where new joiners often come from different countries and need local policy quickly. Third, a foundation for automation: AI customer support, lead qualification and internal assistants all need the same thing, a clean, permissioned, cited knowledge layer. See AI customer support for UAE businesses for the customer-facing use.
A bilingual gap. W3Techs estimates Arabic is the content language of only 0.6% of websites whose language is known (W3Techs). General-purpose AI models therefore see far less Arabic business content than English, which is one more reason to give them your own Arabic material rather than rely on what they learned in training.
What content belongs in a UAE business knowledge base
The answer first: include content that is approved, current, owned and needed for the chosen use case. Start narrow. A customer support knowledge base and an internal HR assistant need different content and different permissions, and should usually be separate indexes.
| Content type | Typical UAE examples | Format challenges | Owner |
|---|---|---|---|
| Customer FAQs and policies | Returns, delivery by emirate, cash-on-delivery rules, warranty | EN and AR versions drift apart | Customer service lead |
| Product or service catalogue | Specifications, sizes, service areas, availability | Changes daily; better read live from the system | Ecommerce or operations |
| Property information | Project brochures, floor plans, service charges, payment plans | Image-heavy PDFs, tables, frequent updates | Sales operations |
| Hospitality information | Facilities, dining hours, Ramadan timings, transfers | Seasonal changes | Front office |
| HR policies | Leave, working hours, visa and Emirates ID processes, onboarding | Restricted access; free zone vs mainland differences | HR |
| Internal SOPs | Order handling, approvals, escalation paths | Often undocumented or in chat | Process owners |
| Service documentation | Installation guides, maintenance manuals, troubleshooting | Scanned manuals, diagrams, mixed languages | Technical lead |
Worth noting
HR policy content should reflect the UAE labour law and any free zone employment rules that apply to each entity. An AI assistant can quote your approved policy; it should not interpret the law. Route legal questions to HR or an adviser.
What should NOT go into the knowledge base
The answer first: if you would not show it to every person who can query the assistant, it does not belong in that index. Retrieval does not respect confidentiality on its own.
- Secrets: passwords, API keys, bank details, contract terms you would not disclose
- Personal data without a lawful basis: customer lists, employee files, CVs, Emirates ID copies. The UAE PDPL requires consent unless an exception applies
- Outdated drafts and superseded versions: keep one current version; archive the rest outside the index
- Unapproved pricing, discounts and offers: read prices live from the system that owns them
- Health data unless hosting, permissions and contracts meet Federal Law No. 2 of 2019 and, in Abu Dhabi, ADHICS
- Legal opinions and board papers in any general-access index
- Content you do not have rights to use, such as third-party reports under restrictive licences
Document processing: PDFs, scanned Arabic and tables
The answer first: retrieval quality is capped by parsing quality. If the text extracted from a document is wrong, out of order or missing its table structure, no embedding model or prompt will fix it.
Digital PDFs and web pages. Extract text with its structure: headings, lists, tables and page numbers. Keep the heading path (for example 'Returns > Electronics > Time limits') as metadata, because it helps both retrieval and citations.
Scanned Arabic documents. Many UAE businesses hold scanned contracts, trade licences, tenancy documents and supplier letters, often with Arabic and English side by side, stamps and signatures. Use an OCR engine that lists Arabic, test it on your own scans, and check the output by eye: right-to-left reading order, joined letters, and digits are common failure points. Low-quality phone photos of documents need more review. See intelligent document processing for extraction methods.
Tables. Price lists, payment plans, service schedules and specification sheets are tables. Convert each table to a structured form (for example Markdown or rows with headers repeated) and keep it in one chunk where possible, so a row is never separated from its column headers.
Bilingual side-by-side layouts. Where a document prints Arabic and English in two columns, parse each column separately and tag the language, otherwise the extracted text interleaves the two languages line by line.
Chunking: starting points, not rules
The answer first: split documents along their natural structure (sections, clauses, Q&A pairs, table boundaries), and use a provider default size as a starting point. Then test with your own questions.
Published defaults. OpenAI's file search defaults to a max chunk size of 800 tokens with 400 tokens of overlap, configurable between 100 and 4,096 tokens, with overlap no more than half the chunk size (OpenAI). Microsoft's Azure AI Search guidance says: 'We recommend starting with a chunk size of 512 tokens (approximately 2,000 characters) and an initial overlap of 25%, which equals 128 tokens' (Microsoft Learn). Two credible providers recommend different numbers, which tells you there is no universal answer.
Arabic note. Token counts differ between languages and tokenisers, and Arabic text often uses more tokens per word than English in many tokenisers. Measure chunk sizes in tokens with the tokeniser your embedding model uses, rather than assuming a character count carries over from English.
Metadata on every chunk. At minimum: document title, section path, language, document owner, effective date, access group and a link to the source. Language and access group are essential for the UAE patterns later in this guide. Our RAG chunking strategies guide compares methods in detail.
Embeddings: multilingual, but verify Arabic
The answer first: use a multilingual embedding model so that Arabic and English text about the same thing land near each other, but confirm Arabic is on the provider's published language list and test it on your content.
What we could verify. Cohere documents its embed-multilingual-v3.0 model as supporting 'over 100 languages' (Cohere); the page we checked did not name Arabic individually, so check the provider's full language list. For other providers we could not verify an official Arabic support statement at the time of writing. Our advice is the same for every vendor: check the provider's language list and test.
How to test. Take 50 questions in Arabic and 50 in English with known correct passages, embed them, and measure how often the correct passage appears in the top results. Repeat for English questions over Arabic documents and the reverse. Compare two or three models on the same set before committing. For background, read vector embeddings explained and vector databases for AI.
Retrieval: vector, keyword, hybrid, reranking and contextual retrieval
The answer first: use hybrid search (keyword plus vector), then rerank the top candidates. This combination is the most reliable default for business content, and it matters more in a bilingual UAE knowledge base full of product codes, project names and Arabic spelling variants.
Vector search finds passages with similar meaning even when the words differ. Keyword search, usually BM25, finds exact terms. Microsoft notes keyword search is better for 'product codes, highly specialized jargon, dates, and people's names'. Azure AI Search runs both 'in parallel' and merges them with Reciprocal Rank Fusion, and Microsoft says benchmark testing indicates hybrid retrieval with semantic ranking 'offers significant benefits in search relevance' (Microsoft Learn). OpenAI's file search likewise retrieves 'through semantic and keyword search'. See hybrid search for RAG.
Reranking takes the top candidates from first-stage retrieval and re-scores them with a model that reads the question and passage together. Cohere describes rerank-v4.0-pro as a multilingual model (Cohere). Read RAG reranking for tuning candidate counts.
Contextual retrieval. A chunk such as 'The fee is waived for renewals' is ambiguous on its own. Anthropic's contextual retrieval method adds a short, generated context of 50 to 100 tokens to each chunk before indexing (for example, which policy and which product it belongs to). In Anthropic's own internal tests, measured as the top-20 retrieval failure rate from a 5.7% baseline, contextual embeddings reduced failures by 35%, contextual embeddings plus contextual BM25 by 49%, and adding reranking by 67%, to 1.9% (Anthropic). Those are Anthropic's benchmark results on its datasets, not a guarantee for yours, but they show the value of combining the techniques.
| Method | Strong at | Weak at | Use it for |
|---|---|---|---|
| Vector search | Paraphrases, cross-lingual meaning | Exact codes, names, numbers | Natural-language questions |
| Keyword (BM25) | Exact terms, SKUs, project names, dates | Synonyms, other languages | Identifiers and jargon |
| Hybrid | Both of the above | Needs tuning of fusion | The default |
| Reranking | Precision in the final few results | Adds latency and cost | After hybrid retrieval |
| Contextual chunks | Ambiguous, context-dependent passages | Extra processing at ingestion | Policies, contracts, manuals |
Generation, citations, refusal and human escalation
The answer first: the model should answer only from retrieved passages, cite each claim, say clearly when the sources do not answer the question, and hand over to a person in those cases.
Citations. Anthropic's Citations feature is generally available and, in Anthropic's description, returns 'the exact passages that support each claim' (Anthropic). Other platforms offer similar features. In the interface, show the document title, section and date, and link to the source so staff and customers can check.
Refusal. Write the rule plainly: if the retrieved passages do not contain the answer, the assistant says so and offers the next step. Test it with questions that are deliberately outside the content. A knowledge base assistant that never says 'I don't know' is guessing some of the time. See reducing hallucinations in AI agents.
Escalation. For customer-facing use, unanswered questions go to a support queue with the question, language and retrieved passages attached; for internal use, to the content owner. Each unanswered question is a candidate for new content.
Injected instructions. Documents and web pages can contain text that tries to instruct the model. Treat retrieved content as data, never as instructions. Read indirect prompt injection.
A practical architecture
The answer first: two pipelines share one index. The ingestion pipeline turns approved documents into permissioned, language-tagged chunks. The query pipeline retrieves, reranks and generates a cited answer or escalates. The diagram is our recommended reference design; enterprise RAG architecture covers scaling it across departments.
| Component | Purpose | Options to consider |
|---|---|---|
| Connectors | Pull content from where it lives | SharePoint, Google Drive, CMS, helpdesk, database exports |
| Parser and OCR | Extract text, structure and tables | Document AI services; test Arabic scans specifically |
| Normaliser | Make Arabic spelling variants match | Lucene or Elasticsearch Arabic analysers, or equivalent |
| Chunker | Split along structure; add context | Heading-aware splitting; contextual prefixes |
| Embedding model | Represent meaning across languages | A multilingual model whose list includes Arabic |
| Index | Store vectors, text and metadata | A search service with hybrid search and filters; see vector databases |
| Reranker | Improve top results | A multilingual reranking model |
| Language model | Write the cited answer | Choose by quality in both languages, region and terms |
| Access control | Filter results by user | Security filters on group IDs; synced permissions |
| Evaluation and logs | Measure and improve | Bilingual test set; sampled review; feedback buttons |
INGESTION
Sources (SharePoint, Drive, CMS, PDFs, scans, DB)
-> Parse (text, headings, tables) + OCR for scans
-> Clean + tag (language, owner, date, access group)
-> Arabic normalisation (for keyword index)
-> Chunk (structure-aware) + context prefix
-> Embed (multilingual model)
-> Index: vector + keyword, with metadata
QUERY
User question (EN / AR / mixed)
-> Identify user + permissions
-> Detect language, normalise query
-> Retrieve: hybrid search, filtered by access
-> Rerank top candidates
-> Generate answer from passages, with citations
-> Sources sufficient?
yes -> answer + citations
no -> say so + escalate to a person
-> Log question, sources, answer, feedbackArabic and English knowledge bases: what changes
The answer first: a bilingual knowledge base is not two monolingual ones side by side. You need Arabic text normalisation, cross-lingual retrieval, a rule for which version is authoritative, a process to keep versions in sync, attention to Arabic OCR, and evaluation in both languages.
Normalisation. Arabic has several spellings for what users treat as the same word. Lucene's ArabicNormalizer, which Elasticsearch and OpenSearch use in their Arabic analysis, normalises hamza forms of alef to a bare alef (أ إ آ → ا), teh marbuta to heh (ة → ه), alef maksura to yeh (ى → ي), and removes diacritics (harakat) and tatweel, the stretching character (Apache Lucene). Elasticsearch's Arabic analyser also adds stop words, digit folding and stemming (Elastic). Apply the same normalisation to documents and queries in the keyword index. Diacritics matter here because Modern Standard Arabic is typically written without them, so a vowelled document and an unvowelled query must still match.
Cross-lingual retrieval. A customer may ask in English about a policy that exists only in Arabic, or the reverse. Multilingual embeddings can match across languages, but keyword search cannot. Options: index both language versions where they exist; translate the query and search in both languages; or store a reviewed translation of key documents. Test English-over-Arabic and Arabic-over-English questions explicitly.
Dialect. Customers write in Gulf, Levantine, Egyptian and other dialects; most business documents are in MSA. Academic work has found that a query in MSA may not retrieve colloquial Arabic documents, and the reverse is likely too. Include dialect questions in your test set and add common dialect terms to synonyms where retrieval misses them.
Parallel versions and authority. Decide, per document type, which language is authoritative. For consumer-facing content this may be Arabic: UAE consumer protection rules require consumer invoices in Arabic and product or service information in Arabic for UAE-registered ecommerce businesses (u.ae). Many internal SOPs are written in English first. Record the authoritative language in metadata, link each translation to its source with a version number, and when the source changes, mark the translation stale until a fluent reviewer updates it. If the assistant finds conflicting versions, it should prefer the authoritative one and flag the conflict.
Arabic OCR quality. Treat OCR output from Arabic scans as untrusted until sampled. Track an error log by document source, and re-scan or retype high-value documents rather than indexing poor text.
Evaluation in both languages. Build separate Arabic and English test sets, plus a cross-lingual set, and report results separately. A combined score can hide a weak Arabic experience. For how this connects to your public site's Arabic content, see Arabic SEO for UAE businesses and multilingual website development in the UAE.
| Normalisation step | Example | Why it matters |
|---|---|---|
| Alef with hamza → bare alef | أ / إ / آ → ا | Users type hamza inconsistently |
| Teh marbuta → heh | ة → ه | Word endings are often typed either way |
| Alef maksura → yeh | ى → ي | Common variation, especially in Gulf typing |
| Remove diacritics | Strip harakat | Most text is unvowelled; some documents are not |
| Remove tatweel | Strip the stretching character ـ | Used decoratively in headings and brochures |
Access control and security trimming
The answer first: a user must never receive an answer built from a document they are not allowed to open. Enforce this in the retrieval layer, before passages reach the model, using the same permissions as the source system.
How it is done. Microsoft documents several approaches for Azure AI Search, which it calls essential for RAG and agentic systems: security filters (generally available), POSIX-like ACL and RBAC scopes, Microsoft Purview sensitivity labels and SharePoint ACLs (the last three in preview at the time of checking) (Microsoft Learn). The security filter pattern 'trims search results based on a string containing a group or user identity' (Microsoft Learn). Other search engines support equivalent metadata filters.
Practical rules. Store an access group on every chunk at ingestion. Sync permissions when they change in the source system, and remove deleted documents from the index promptly. Keep separate indexes for clearly separate audiences (customers, all staff, HR, management). Log which documents were used in each answer. See AI agent access control.
Data protection and residency in the UAE
UAE facts. The Personal Data Protection Law (Federal Decree-Law No. 45 of 2021) has been in force since 2 January 2022; consent is required unless an exception applies, and cross-border transfer conditions apply. We could not find officially published executive regulations as of October 2026 (u.ae). The DIFC and ADGM have their own data protection regimes. Health data: Article 13 of Federal Law No. 2 of 2019 restricts storing or processing health data outside the UAE, and Abu Dhabi's ADHICS standard requires UAE hosting, including backup and disaster recovery, for in-scope health information.
In-country processing options (facts only). AWS operates a UAE region (me-central-1, launched 2022). Microsoft Azure has UAE North (Dubai) and UAE Central (Abu Dhabi, restricted). Oracle has Dubai and Abu Dhabi regions. Google Cloud has no UAE region; its nearest are Doha and Dammam. Check the region of every component, including the language model endpoint, embeddings, OCR, logs and backups, not only the index.
Our recommendation. Classify content before ingestion (public, internal, confidential, personal, health), decide the hosting and model options each class allows, and record the decision. Read your AI vendors' data-use terms on training and retention. This is not legal advice: confirm your obligations with an adviser, especially for health, financial and DIFC or ADGM entities. See AI data privacy.
UAE examples
These are hypothetical examples showing typical content, users and controls. They are not ZSpace clients.
| Use case | Content | Users | Key controls |
|---|---|---|---|
| Property information | Project brochures, floor plans, payment plans, service charge notes, FAQs in EN and AR | Sales agents, website visitors | Prices read live or dated; no investment or legal advice |
| Product catalogue | Specifications, compatibility, warranty, care guides | Customers on WhatsApp and web, support staff | Stock and price from the ecommerce system, not documents |
| HR policies | Leave, working hours, visa and onboarding steps per entity | Employees | Access by entity and role; legal questions to HR |
| Customer support | Returns, delivery, payment and complaints policies | Support agents first, customers later | Citations, refusal rule, escalation queue |
| Hospitality information | Facilities, dining, transfers, seasonal timings | Guests, front office | Owners for seasonal content; expiry dates |
| Internal SOPs | Approvals, order handling, incident steps | Operations staff | Version control; one current SOP per process |
| Service documentation | Manuals, troubleshooting, maintenance checklists | Field technicians | Scanned manuals checked; offline access needs |
Step-by-step implementation roadmap
This is our recommended sequence. Most of the effort goes into content and evaluation, not code. If you are assessing broader AI readiness first, see agentic AI readiness for UAE businesses.
| Step | What to do | Output |
|---|---|---|
| 1. Choose the use case | One audience, one job (for example, support agents answering policy questions) | Scope, success measures, owner |
| 2. Inventory content | List sources, owners, languages, dates and access groups | Content register |
| 3. Clean and approve | Remove stale drafts; fix EN/AR mismatches; mark authoritative language | Approved corpus |
| 4. Build the test set | 100–300 real questions in EN, AR and cross-lingual with correct sources | Bilingual evaluation set |
| 5. Ingest | Parse, OCR, normalise, chunk, tag metadata, embed, index | Searchable index with permissions |
| 6. Tune retrieval | Compare chunking, hybrid weights, reranking on the test set | Measured retrieval baseline |
| 7. Generate with citations | Grounding rules, refusal rule, escalation path | Assistant ready for internal pilot |
| 8. Pilot internally | Staff use it; sample answers reviewed in both languages | Accuracy findings, content backlog |
| 9. Launch and monitor | Release to the target audience; feedback buttons; weekly review | Live service with owners |
| 10. Maintain | Review dates, permission sync, re-run tests on every change | Stable quality over time |
Evaluation without fabricated benchmarks
The answer first: vendor benchmarks do not tell you how a system will perform on your documents and your customers' Arabic. Build your own test set, measure a few clear metrics, and re-run them after every change.
Building the set. Collect real questions from support tickets, WhatsApp chats, staff and search logs. For each, record the correct answer, the source passage, the language and whether the question should be refused. Include tricky cases: dialect, Arabizi, product codes, outdated policies and questions outside the content. Keep the set private so it is not used to tune prompts directly. Our LLM evaluation pipeline guide covers automation.
| Metric | Definition | How to measure |
|---|---|---|
| Retrieval hit rate | Share of questions where the correct passage is in the top k results | Automatic, from the labelled test set |
| Answer faithfulness | Share of answers fully supported by the retrieved passages | Human review, or a model-graded check that humans spot-check |
| Answer correctness | Share of answers that match the approved answer | Human review against the reference |
| Citation accuracy | Share of citations pointing to a passage that supports the claim | Human review of a sample |
| Correct refusal rate | Share of out-of-scope questions the assistant declines | Automatic, from labelled refusal cases |
| Language parity | Gap between Arabic and English scores | Compare metrics per language |
| Freshness | Share of documents past their review date | Content register report |
Common mistakes
Indexing everything. A shared drive full of drafts produces confident, outdated answers.
English-only testing. The Arabic experience is discovered by customers, not by the team.
No authoritative language. Two versions of a policy disagree and the assistant picks one at random.
Skipping Arabic normalisation. Keyword search misses documents because of a hamza or a ta marbuta.
Vector search only. Product codes, project names and dates are missed.
Permissions in the prompt. Telling the model not to reveal a document is not access control.
Copying prices into documents. They go stale; read them from the system of record.
No owner after launch. Content ages, permissions drift and quality falls quietly.
Ignoring where the model runs. The index is in the UAE, but the model endpoint, OCR or logs are not.
Sources
Retrieval and models: OpenAI, retrieval and file search chunking; OpenAI, file search tool; Microsoft, chunking documents; Microsoft, hybrid search; Microsoft, document-level access; Microsoft, security trimming; Anthropic, contextual retrieval; Anthropic, Citations; Cohere Embed; Cohere Rerank.
Arabic text processing: Apache Lucene ArabicNormalizer; Elasticsearch language analysers; W3Techs content languages.
UAE: u.ae data protection laws; u.ae consumer protection; Microsoft AI Economy Institute; AWS and UAE AI Office study.
Benchmark figures are attributed to the organisations that published them and are not ZSpace data. Regulations change: confirm data protection and health data obligations with the relevant authority or an adviser.
Conclusion
A useful AI knowledge base for a UAE business is mostly a content and governance project with a retrieval system attached. Pick one use case, approve and date the content, decide which language is authoritative, normalise Arabic, use hybrid search with reranking, cite every answer, enforce permissions in retrieval, keep data where your obligations require it, and measure with your own bilingual test set. Done that way, the same knowledge layer can serve support staff, customers on WhatsApp and internal teams. When you are ready to put it in front of customers, our guide to AI customer support for UAE businesses covers the channels and escalation.
Planning a bilingual knowledge base?
ZSpace Labs is an India-based, remote-first technology studio working with UAE and global businesses on AI and workflow automation and web platforms. We can help inventory and clean your content, design Arabic and English retrieval, and build an evaluation set before anything goes live.
Common questions.
An AI knowledge base is a curated, searchable collection of a business's approved documents and data that an AI assistant retrieves from before answering, so its answers are grounded in your content and can cite their sources. It is usually built with retrieval-augmented generation (RAG): documents are parsed, split into chunks, indexed, retrieved for each question and passed to a language model.