← back to the directory
Model Inference & Serving

LLaVA-NeXT

1Open 2vision-language 3assistant 4built 5on 6Llama

the six Ws · specification

W1 Who

Researchers at University of Wisconsin-Madison, Microsoft Research, and Columbia University created LLaVA-NeXT.

W2 What

An open multimodal model that combines a vision encoder with a Llama or Vicuna language model for image-grounded chat.

W3 Where

Weights and code are hosted on Hugging Face and GitHub under the LLaVA project.

W4 When

The LLaVA-NeXT (1.6) update was released in January 2024, following the original LLaVA in 2023.

W5 Why

It demonstrated that visual instruction tuning could produce GPT-4V-level multimodal chat at low training cost.

W6 With

Runs via the LLaVA/llava-next codebase on PyTorch and transformers, typically served through vLLM or Ollama.

W7 Watch

Open weights released by the LLaVA research team; license follows the underlying Llama/Vicuna base model terms, permitting research and most commercial use with restrictions.

vision-languagemultimodalopen-sourcevisual-instruction-tuning

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.