Languages supported by Content Understanding

  • Release version: Australia
  • Updated March 12, 2026
  • 1 minute to read
  • Use this reference to identify which languages the Content Understanding application supports for text-based and image-based files, and which optical character recognition (OCR) model applies to each language group.

    For text-based files, the AI recognizes any language supported by the selected or default model, as described in the model card for the LLM. For more information on LLMs, see Large language models used by Content Understanding.

    For image files that require OCR to detect text, OCR models support different language groups. The language selected during use case setup helps the OCR model detect text in the images. For more information, see Set up a use case.

    The following OCR models are available.

    Table 1. OCR models for Content Understanding
    OCR model Languages supported
    Latin model (Default) English, Dutch, French, German, Italian, Portuguese, and Spanish
    CJ model Japanese and Chinese (simplified)