Command A Vision

Available · Language, Multimodal

Command A Vision is a 112B-parameter model in Cohere's Command A family that reads images alongside text and answers in text, released on 31 July 2025 on the Cohere API with open weights for non-commercial research use. Cohere calls it its first commercial model that can interpret images and aims it at charts, diagrams, tables, document OCR and scene analysis. [1] [2] [3] Primary source

Timeline of Command A Vision →

Claims and evidence

  • Released Primary source[2]page date; Availability [1]entry date
  • Status Available Primary source[4]Command table: command-a-vision-07-2025, Status Live [3]
  • Derived from (fine tune) Command AThe model card describes a language model based on Command A paired with a vision encoder through a multimodal adapter. The method fine-tune follows the Hugging Face model tree classification of the base_model metadata; the card text does not name the training method. Primary source[3]card metadata base_model CohereLabs/c4ai-command-a-03-2025; Model Details: Model Architecture
  • Change · Modality Adds image input next to text, up to 20 images per request; Command A accepts text only.Compared with Command A Primary source[4]Command table, Modality [1]
  • Change · Context length Context window of 128K tokens, against 256K for Command A.Compared with Command A Primary source[4]Command table, Context Length
  • Change · Tool use Tool use is not supported with this model, according to Cohere's model documentation.Compared with Command A Primary source[5]Limitations
  • Change · Languages Officially supports six languages: English, Portuguese, Italian, French, German and Spanish.Compared with Command A Primary source[1]
  • Input text, image Primary source[4]Command table, Modality [3]Model Details: Input
  • Output text Primary source[3]Model Details: Output [5]Limitations
  • Open weights YesResearch release under CC-BY-NC 4.0 with Cohere Labs acceptable use policy; gated download on Hugging Face. Primary source[3]Model Summary; Terms of Use [2]Availability
  • Context window 128K tokensMaximum output 8K tokens per the changelog and the models overview. The Hugging Face card says the model supports 128K but is configured there for 32K. Primary source[4]Command table, Context Length [1]Technical Specifications [5]Description
  • Parameters 112B Primary source[3]Model Summary: Model Size
  • Access API, Open-weights download Primary source[2]Availability [1]Technical Specifications: API Endpoint

Lineage

Predecessors

No known predecessor.

Successors

No known successor.

Based on

  • Command A · 13 March 2025 · derived (fine tune)

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Cohere on AI Radar →