Skip to content
AI & Automation20 min read

How to Build an AI Knowledge Base for a UAE Business

How to build an AI knowledge base for a UAE business: bilingual Arabic and English content, chunking, hybrid retrieval, citations, access control and residency.

01

Quick answer

An AI knowledge base is your business's approved documents and data, prepared so an AI assistant can find the right passage and answer from it with a citation. For a UAE business, build it in six steps: choose a use case, collect and clean bilingual content, parse and chunk it, index it for hybrid search, generate cited answers, and escalate when sources are missing.

The UAE-specific work is in the detail: Arabic and English versions of the same policy, scanned Arabic documents, Arabic spelling variants, data protection rules that differ between the mainland, the DIFC and ADGM, and health data that must stay in the country. This guide focuses on those points and links to our RAG guide and the rest of the retrieval cluster for technical depth. For a shorter, general introduction, read AI knowledge base: building an assistant on company documents.

02

Key takeaways

  • Start with one use case and one audience; the content and permissions follow from that.
  • Leave out secrets, unapproved prices, stale drafts and personal data without a lawful basis.
  • Use provider chunking defaults as starting points (800/400 tokens at OpenAI, 512 tokens with 25% overlap in Azure guidance), then test.
  • Combine keyword and vector search; keyword search catches product codes, names and dates that vectors miss.
  • Normalise Arabic (alef forms, ta marbuta, alef maksura, diacritics, tatweel) so spelling variants match.
  • Decide which language version is authoritative, and keep parallel versions in sync.
  • Enforce document-level permissions in retrieval, not in the prompt.
  • Evaluate with your own bilingual test set; do not rely on vendor benchmarks.
03

What an AI knowledge base is, in exact terms

Definition: an AI knowledge base is a governed collection of approved content (documents, pages, records and structured data), converted into a searchable index so that an AI system can retrieve the most relevant passages for a question and generate an answer grounded in them. The pattern behind it is retrieval-augmented generation (RAG).

It differs from a traditional help centre, which customers search themselves, and from a chatbot with scripted answers. It is also different from training a model: the model does not memorise your documents; it reads the retrieved passages at answer time. That is why it can stay current and cite sources. For when training does make sense, see RAG vs fine-tuning.

TermDefinition
IngestionCollecting source content and converting it into clean text with structure and metadata
Parsing / OCRExtracting text, headings and tables from files; OCR (optical character recognition) reads text from scanned images
ChunkA passage of a document, sized for retrieval, stored with metadata such as title, language, owner and date
EmbeddingA list of numbers representing a chunk's meaning, so similar meanings sit close together
Vector searchFinding chunks whose embeddings are closest to the question's embedding
Keyword (BM25) searchRanking chunks by matching words, weighted by how rare and frequent they are
Hybrid searchRunning keyword and vector search together and merging the results
RerankingA second model re-scoring the top candidates for relevance to the question
Grounded generationA model writing the answer only from retrieved passages
CitationA pointer from a claim in the answer to the passage that supports it
Security trimmingRemoving results the user is not allowed to see before they reach the model
04

Why UAE businesses need one

UAE context. Generative AI use is very high: Microsoft's AI Economy Institute estimates 70.1% of the UAE working-age population used a generative AI product in Q1 2026 (Microsoft). An AWS and UAE AI Office study reports 72% of UAE businesses have adopted AI (Zawya). When staff and customers ask AI tools about your business and the tools have no access to your approved content, they guess, or staff paste company documents into consumer tools.

Three practical reasons. First, consistent answers: support, sales and operations staff answer from the same approved source, in English and Arabic. Second, faster onboarding in a workforce where new joiners often come from different countries and need local policy quickly. Third, a foundation for automation: AI customer support, lead qualification and internal assistants all need the same thing, a clean, permissioned, cited knowledge layer. See AI customer support for UAE businesses for the customer-facing use.

A bilingual gap. W3Techs estimates Arabic is the content language of only 0.6% of websites whose language is known (W3Techs). General-purpose AI models therefore see far less Arabic business content than English, which is one more reason to give them your own Arabic material rather than rely on what they learned in training.

05

What content belongs in a UAE business knowledge base

The answer first: include content that is approved, current, owned and needed for the chosen use case. Start narrow. A customer support knowledge base and an internal HR assistant need different content and different permissions, and should usually be separate indexes.

Content typeTypical UAE examplesFormat challengesOwner
Customer FAQs and policiesReturns, delivery by emirate, cash-on-delivery rules, warrantyEN and AR versions drift apartCustomer service lead
Product or service catalogueSpecifications, sizes, service areas, availabilityChanges daily; better read live from the systemEcommerce or operations
Property informationProject brochures, floor plans, service charges, payment plansImage-heavy PDFs, tables, frequent updatesSales operations
Hospitality informationFacilities, dining hours, Ramadan timings, transfersSeasonal changesFront office
HR policiesLeave, working hours, visa and Emirates ID processes, onboardingRestricted access; free zone vs mainland differencesHR
Internal SOPsOrder handling, approvals, escalation pathsOften undocumented or in chatProcess owners
Service documentationInstallation guides, maintenance manuals, troubleshootingScanned manuals, diagrams, mixed languagesTechnical lead

Worth noting

HR policy content should reflect the UAE labour law and any free zone employment rules that apply to each entity. An AI assistant can quote your approved policy; it should not interpret the law. Route legal questions to HR or an adviser.

06

What should NOT go into the knowledge base

The answer first: if you would not show it to every person who can query the assistant, it does not belong in that index. Retrieval does not respect confidentiality on its own.

  • Secrets: passwords, API keys, bank details, contract terms you would not disclose
  • Personal data without a lawful basis: customer lists, employee files, CVs, Emirates ID copies. The UAE PDPL requires consent unless an exception applies
  • Outdated drafts and superseded versions: keep one current version; archive the rest outside the index
  • Unapproved pricing, discounts and offers: read prices live from the system that owns them
  • Health data unless hosting, permissions and contracts meet Federal Law No. 2 of 2019 and, in Abu Dhabi, ADHICS
  • Legal opinions and board papers in any general-access index
  • Content you do not have rights to use, such as third-party reports under restrictive licences
07

Document processing: PDFs, scanned Arabic and tables

The answer first: retrieval quality is capped by parsing quality. If the text extracted from a document is wrong, out of order or missing its table structure, no embedding model or prompt will fix it.

Digital PDFs and web pages. Extract text with its structure: headings, lists, tables and page numbers. Keep the heading path (for example 'Returns > Electronics > Time limits') as metadata, because it helps both retrieval and citations.

Scanned Arabic documents. Many UAE businesses hold scanned contracts, trade licences, tenancy documents and supplier letters, often with Arabic and English side by side, stamps and signatures. Use an OCR engine that lists Arabic, test it on your own scans, and check the output by eye: right-to-left reading order, joined letters, and digits are common failure points. Low-quality phone photos of documents need more review. See intelligent document processing for extraction methods.

Tables. Price lists, payment plans, service schedules and specification sheets are tables. Convert each table to a structured form (for example Markdown or rows with headers repeated) and keep it in one chunk where possible, so a row is never separated from its column headers.

Bilingual side-by-side layouts. Where a document prints Arabic and English in two columns, parse each column separately and tag the language, otherwise the extracted text interleaves the two languages line by line.

08

Chunking: starting points, not rules

The answer first: split documents along their natural structure (sections, clauses, Q&A pairs, table boundaries), and use a provider default size as a starting point. Then test with your own questions.

Published defaults. OpenAI's file search defaults to a max chunk size of 800 tokens with 400 tokens of overlap, configurable between 100 and 4,096 tokens, with overlap no more than half the chunk size (OpenAI). Microsoft's Azure AI Search guidance says: 'We recommend starting with a chunk size of 512 tokens (approximately 2,000 characters) and an initial overlap of 25%, which equals 128 tokens' (Microsoft Learn). Two credible providers recommend different numbers, which tells you there is no universal answer.

Arabic note. Token counts differ between languages and tokenisers, and Arabic text often uses more tokens per word than English in many tokenisers. Measure chunk sizes in tokens with the tokeniser your embedding model uses, rather than assuming a character count carries over from English.

Metadata on every chunk. At minimum: document title, section path, language, document owner, effective date, access group and a link to the source. Language and access group are essential for the UAE patterns later in this guide. Our RAG chunking strategies guide compares methods in detail.

09

Embeddings: multilingual, but verify Arabic

The answer first: use a multilingual embedding model so that Arabic and English text about the same thing land near each other, but confirm Arabic is on the provider's published language list and test it on your content.

What we could verify. Cohere documents its embed-multilingual-v3.0 model as supporting 'over 100 languages' (Cohere); the page we checked did not name Arabic individually, so check the provider's full language list. For other providers we could not verify an official Arabic support statement at the time of writing. Our advice is the same for every vendor: check the provider's language list and test.

How to test. Take 50 questions in Arabic and 50 in English with known correct passages, embed them, and measure how often the correct passage appears in the top results. Repeat for English questions over Arabic documents and the reverse. Compare two or three models on the same set before committing. For background, read vector embeddings explained and vector databases for AI.

10

Retrieval: vector, keyword, hybrid, reranking and contextual retrieval

The answer first: use hybrid search (keyword plus vector), then rerank the top candidates. This combination is the most reliable default for business content, and it matters more in a bilingual UAE knowledge base full of product codes, project names and Arabic spelling variants.

Vector search finds passages with similar meaning even when the words differ. Keyword search, usually BM25, finds exact terms. Microsoft notes keyword search is better for 'product codes, highly specialized jargon, dates, and people's names'. Azure AI Search runs both 'in parallel' and merges them with Reciprocal Rank Fusion, and Microsoft says benchmark testing indicates hybrid retrieval with semantic ranking 'offers significant benefits in search relevance' (Microsoft Learn). OpenAI's file search likewise retrieves 'through semantic and keyword search'. See hybrid search for RAG.

Reranking takes the top candidates from first-stage retrieval and re-scores them with a model that reads the question and passage together. Cohere describes rerank-v4.0-pro as a multilingual model (Cohere). Read RAG reranking for tuning candidate counts.

Contextual retrieval. A chunk such as 'The fee is waived for renewals' is ambiguous on its own. Anthropic's contextual retrieval method adds a short, generated context of 50 to 100 tokens to each chunk before indexing (for example, which policy and which product it belongs to). In Anthropic's own internal tests, measured as the top-20 retrieval failure rate from a 5.7% baseline, contextual embeddings reduced failures by 35%, contextual embeddings plus contextual BM25 by 49%, and adding reranking by 67%, to 1.9% (Anthropic). Those are Anthropic's benchmark results on its datasets, not a guarantee for yours, but they show the value of combining the techniques.

MethodStrong atWeak atUse it for
Vector searchParaphrases, cross-lingual meaningExact codes, names, numbersNatural-language questions
Keyword (BM25)Exact terms, SKUs, project names, datesSynonyms, other languagesIdentifiers and jargon
HybridBoth of the aboveNeeds tuning of fusionThe default
RerankingPrecision in the final few resultsAdds latency and costAfter hybrid retrieval
Contextual chunksAmbiguous, context-dependent passagesExtra processing at ingestionPolicies, contracts, manuals
11

Generation, citations, refusal and human escalation

The answer first: the model should answer only from retrieved passages, cite each claim, say clearly when the sources do not answer the question, and hand over to a person in those cases.

Citations. Anthropic's Citations feature is generally available and, in Anthropic's description, returns 'the exact passages that support each claim' (Anthropic). Other platforms offer similar features. In the interface, show the document title, section and date, and link to the source so staff and customers can check.

Refusal. Write the rule plainly: if the retrieved passages do not contain the answer, the assistant says so and offers the next step. Test it with questions that are deliberately outside the content. A knowledge base assistant that never says 'I don't know' is guessing some of the time. See reducing hallucinations in AI agents.

Escalation. For customer-facing use, unanswered questions go to a support queue with the question, language and retrieved passages attached; for internal use, to the content owner. Each unanswered question is a candidate for new content.

Injected instructions. Documents and web pages can contain text that tries to instruct the model. Treat retrieved content as data, never as instructions. Read indirect prompt injection.

12

A practical architecture

The answer first: two pipelines share one index. The ingestion pipeline turns approved documents into permissioned, language-tagged chunks. The query pipeline retrieves, reranks and generates a cited answer or escalates. The diagram is our recommended reference design; enterprise RAG architecture covers scaling it across departments.

ComponentPurposeOptions to consider
ConnectorsPull content from where it livesSharePoint, Google Drive, CMS, helpdesk, database exports
Parser and OCRExtract text, structure and tablesDocument AI services; test Arabic scans specifically
NormaliserMake Arabic spelling variants matchLucene or Elasticsearch Arabic analysers, or equivalent
ChunkerSplit along structure; add contextHeading-aware splitting; contextual prefixes
Embedding modelRepresent meaning across languagesA multilingual model whose list includes Arabic
IndexStore vectors, text and metadataA search service with hybrid search and filters; see vector databases
RerankerImprove top resultsA multilingual reranking model
Language modelWrite the cited answerChoose by quality in both languages, region and terms
Access controlFilter results by userSecurity filters on group IDs; synced permissions
Evaluation and logsMeasure and improveBilingual test set; sampled review; feedback buttons
Reference architecture: bilingual AI knowledge base
INGESTION
Sources (SharePoint, Drive, CMS, PDFs, scans, DB)
  -> Parse (text, headings, tables) + OCR for scans
  -> Clean + tag (language, owner, date, access group)
  -> Arabic normalisation (for keyword index)
  -> Chunk (structure-aware) + context prefix
  -> Embed (multilingual model)
  -> Index: vector + keyword, with metadata

QUERY
User question (EN / AR / mixed)
  -> Identify user + permissions
  -> Detect language, normalise query
  -> Retrieve: hybrid search, filtered by access
  -> Rerank top candidates
  -> Generate answer from passages, with citations
  -> Sources sufficient?
       yes -> answer + citations
       no  -> say so + escalate to a person
  -> Log question, sources, answer, feedback
13

Arabic and English knowledge bases: what changes

The answer first: a bilingual knowledge base is not two monolingual ones side by side. You need Arabic text normalisation, cross-lingual retrieval, a rule for which version is authoritative, a process to keep versions in sync, attention to Arabic OCR, and evaluation in both languages.

Normalisation. Arabic has several spellings for what users treat as the same word. Lucene's ArabicNormalizer, which Elasticsearch and OpenSearch use in their Arabic analysis, normalises hamza forms of alef to a bare alef (أ إ آ → ا), teh marbuta to heh (ة → ه), alef maksura to yeh (ى → ي), and removes diacritics (harakat) and tatweel, the stretching character (Apache Lucene). Elasticsearch's Arabic analyser also adds stop words, digit folding and stemming (Elastic). Apply the same normalisation to documents and queries in the keyword index. Diacritics matter here because Modern Standard Arabic is typically written without them, so a vowelled document and an unvowelled query must still match.

Cross-lingual retrieval. A customer may ask in English about a policy that exists only in Arabic, or the reverse. Multilingual embeddings can match across languages, but keyword search cannot. Options: index both language versions where they exist; translate the query and search in both languages; or store a reviewed translation of key documents. Test English-over-Arabic and Arabic-over-English questions explicitly.

Dialect. Customers write in Gulf, Levantine, Egyptian and other dialects; most business documents are in MSA. Academic work has found that a query in MSA may not retrieve colloquial Arabic documents, and the reverse is likely too. Include dialect questions in your test set and add common dialect terms to synonyms where retrieval misses them.

Parallel versions and authority. Decide, per document type, which language is authoritative. For consumer-facing content this may be Arabic: UAE consumer protection rules require consumer invoices in Arabic and product or service information in Arabic for UAE-registered ecommerce businesses (u.ae). Many internal SOPs are written in English first. Record the authoritative language in metadata, link each translation to its source with a version number, and when the source changes, mark the translation stale until a fluent reviewer updates it. If the assistant finds conflicting versions, it should prefer the authoritative one and flag the conflict.

Arabic OCR quality. Treat OCR output from Arabic scans as untrusted until sampled. Track an error log by document source, and re-scan or retype high-value documents rather than indexing poor text.

Evaluation in both languages. Build separate Arabic and English test sets, plus a cross-lingual set, and report results separately. A combined score can hide a weak Arabic experience. For how this connects to your public site's Arabic content, see Arabic SEO for UAE businesses and multilingual website development in the UAE.

Normalisation stepExampleWhy it matters
Alef with hamza → bare alefأ / إ / آ → اUsers type hamza inconsistently
Teh marbuta → hehة → هWord endings are often typed either way
Alef maksura → yehى → يCommon variation, especially in Gulf typing
Remove diacriticsStrip harakatMost text is unvowelled; some documents are not
Remove tatweelStrip the stretching character ـUsed decoratively in headings and brochures
14

Access control and security trimming

The answer first: a user must never receive an answer built from a document they are not allowed to open. Enforce this in the retrieval layer, before passages reach the model, using the same permissions as the source system.

How it is done. Microsoft documents several approaches for Azure AI Search, which it calls essential for RAG and agentic systems: security filters (generally available), POSIX-like ACL and RBAC scopes, Microsoft Purview sensitivity labels and SharePoint ACLs (the last three in preview at the time of checking) (Microsoft Learn). The security filter pattern 'trims search results based on a string containing a group or user identity' (Microsoft Learn). Other search engines support equivalent metadata filters.

Practical rules. Store an access group on every chunk at ingestion. Sync permissions when they change in the source system, and remove deleted documents from the index promptly. Keep separate indexes for clearly separate audiences (customers, all staff, HR, management). Log which documents were used in each answer. See AI agent access control.

15

Data protection and residency in the UAE

UAE facts. The Personal Data Protection Law (Federal Decree-Law No. 45 of 2021) has been in force since 2 January 2022; consent is required unless an exception applies, and cross-border transfer conditions apply. We could not find officially published executive regulations as of October 2026 (u.ae). The DIFC and ADGM have their own data protection regimes. Health data: Article 13 of Federal Law No. 2 of 2019 restricts storing or processing health data outside the UAE, and Abu Dhabi's ADHICS standard requires UAE hosting, including backup and disaster recovery, for in-scope health information.

In-country processing options (facts only). AWS operates a UAE region (me-central-1, launched 2022). Microsoft Azure has UAE North (Dubai) and UAE Central (Abu Dhabi, restricted). Oracle has Dubai and Abu Dhabi regions. Google Cloud has no UAE region; its nearest are Doha and Dammam. Check the region of every component, including the language model endpoint, embeddings, OCR, logs and backups, not only the index.

Our recommendation. Classify content before ingestion (public, internal, confidential, personal, health), decide the hosting and model options each class allows, and record the decision. Read your AI vendors' data-use terms on training and retention. This is not legal advice: confirm your obligations with an adviser, especially for health, financial and DIFC or ADGM entities. See AI data privacy.

16

UAE examples

These are hypothetical examples showing typical content, users and controls. They are not ZSpace clients.

Use caseContentUsersKey controls
Property informationProject brochures, floor plans, payment plans, service charge notes, FAQs in EN and ARSales agents, website visitorsPrices read live or dated; no investment or legal advice
Product catalogueSpecifications, compatibility, warranty, care guidesCustomers on WhatsApp and web, support staffStock and price from the ecommerce system, not documents
HR policiesLeave, working hours, visa and onboarding steps per entityEmployeesAccess by entity and role; legal questions to HR
Customer supportReturns, delivery, payment and complaints policiesSupport agents first, customers laterCitations, refusal rule, escalation queue
Hospitality informationFacilities, dining, transfers, seasonal timingsGuests, front officeOwners for seasonal content; expiry dates
Internal SOPsApprovals, order handling, incident stepsOperations staffVersion control; one current SOP per process
Service documentationManuals, troubleshooting, maintenance checklistsField techniciansScanned manuals checked; offline access needs
17

Step-by-step implementation roadmap

This is our recommended sequence. Most of the effort goes into content and evaluation, not code. If you are assessing broader AI readiness first, see agentic AI readiness for UAE businesses.

StepWhat to doOutput
1. Choose the use caseOne audience, one job (for example, support agents answering policy questions)Scope, success measures, owner
2. Inventory contentList sources, owners, languages, dates and access groupsContent register
3. Clean and approveRemove stale drafts; fix EN/AR mismatches; mark authoritative languageApproved corpus
4. Build the test set100–300 real questions in EN, AR and cross-lingual with correct sourcesBilingual evaluation set
5. IngestParse, OCR, normalise, chunk, tag metadata, embed, indexSearchable index with permissions
6. Tune retrievalCompare chunking, hybrid weights, reranking on the test setMeasured retrieval baseline
7. Generate with citationsGrounding rules, refusal rule, escalation pathAssistant ready for internal pilot
8. Pilot internallyStaff use it; sample answers reviewed in both languagesAccuracy findings, content backlog
9. Launch and monitorRelease to the target audience; feedback buttons; weekly reviewLive service with owners
10. MaintainReview dates, permission sync, re-run tests on every changeStable quality over time
18

Evaluation without fabricated benchmarks

The answer first: vendor benchmarks do not tell you how a system will perform on your documents and your customers' Arabic. Build your own test set, measure a few clear metrics, and re-run them after every change.

Building the set. Collect real questions from support tickets, WhatsApp chats, staff and search logs. For each, record the correct answer, the source passage, the language and whether the question should be refused. Include tricky cases: dialect, Arabizi, product codes, outdated policies and questions outside the content. Keep the set private so it is not used to tune prompts directly. Our LLM evaluation pipeline guide covers automation.

MetricDefinitionHow to measure
Retrieval hit rateShare of questions where the correct passage is in the top k resultsAutomatic, from the labelled test set
Answer faithfulnessShare of answers fully supported by the retrieved passagesHuman review, or a model-graded check that humans spot-check
Answer correctnessShare of answers that match the approved answerHuman review against the reference
Citation accuracyShare of citations pointing to a passage that supports the claimHuman review of a sample
Correct refusal rateShare of out-of-scope questions the assistant declinesAutomatic, from labelled refusal cases
Language parityGap between Arabic and English scoresCompare metrics per language
FreshnessShare of documents past their review dateContent register report
19

Common mistakes

Indexing everything. A shared drive full of drafts produces confident, outdated answers.

English-only testing. The Arabic experience is discovered by customers, not by the team.

No authoritative language. Two versions of a policy disagree and the assistant picks one at random.

Skipping Arabic normalisation. Keyword search misses documents because of a hamza or a ta marbuta.

Vector search only. Product codes, project names and dates are missed.

Permissions in the prompt. Telling the model not to reveal a document is not access control.

Copying prices into documents. They go stale; read them from the system of record.

No owner after launch. Content ages, permissions drift and quality falls quietly.

Ignoring where the model runs. The index is in the UAE, but the model endpoint, OCR or logs are not.

20

Sources

21

Conclusion

A useful AI knowledge base for a UAE business is mostly a content and governance project with a retrieval system attached. Pick one use case, approve and date the content, decide which language is authoritative, normalise Arabic, use hybrid search with reranking, cite every answer, enforce permissions in retrieval, keep data where your obligations require it, and measure with your own bilingual test set. Done that way, the same knowledge layer can serve support staff, customers on WhatsApp and internal teams. When you are ready to put it in front of customers, our guide to AI customer support for UAE businesses covers the channels and escalation.

Planning a bilingual knowledge base?

ZSpace Labs is an India-based, remote-first technology studio working with UAE and global businesses on AI and workflow automation and web platforms. We can help inventory and clean your content, design Arabic and English retrieval, and build an evaluation set before anything goes live.

Start a Project
FAQ

Common questions.

An AI knowledge base is a curated, searchable collection of a business's approved documents and data that an AI assistant retrieves from before answering, so its answers are grounded in your content and can cite their sources. It is usually built with retrieval-augmented generation (RAG): documents are parsed, split into chunks, indexed, retrieved for each question and passed to a language model.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.