Languages supported by Content Understanding
Use this reference to identify which languages the Content Understanding application supports for text-based and image-based files, and which optical character recognition (OCR) model applies to each language group.
For text-based files, the AI recognizes any language supported by the selected or default model, as described in the model card for the LLM. For more information on LLMs, see Large language models used by Content Understanding.
For image files that require OCR to detect text, OCR models support different language groups. The language selected during use case setup helps the OCR model detect text in the images. For more information, see Set up a use case.
The following OCR models are available.
| OCR model | Languages supported |
|---|---|
| Latin model (Default) | English, Dutch, French, German, Italian, Portuguese, and Spanish |
| CJ model | Japanese and Chinese (simplified) |