Vision–language–action model
Foundation model allowing control of robot actions
In robot learning, a vision–language–action model (VLA) is a class of multimodal foundation models that integrates vision, language and actions. Given an input image (or video) of the robot's surroundings and a text instruction, a VLA directly outputs low-level robot actions that can be executed to accomplish the requested task.
Nº Q133568269 ★★
Poco común · Saberes
Vision–language–action model
Foundation model allowing control of robot actions
In robot learning, a vision–language–action model (VLA) is a class of multimodal foundation models that integrates vision, language and actions. Given an input image (or video) of the robot's surroundings and a text instruction, a VLA directly outputs low-level robot actions that can be executed to accomplish the requested task.
Último precio
—
Precio mínimo
—
Mediana 7 d
—
Ventas 30 d
0
Rango 30 d
—
En circulación
0
Cotización
mediana
mín – máx
ventas
Sin ventas en el periodo
Ver tabla
| Fecha | mediana | Mín | Máx | ventas |
|---|
Historial de ventas
- Última venta
- —
- Media 30 d
- —
- Mínimo 30 d
- —
- Máximo 30 d
- —
- Ventas 7 d
- 0
- Ventas 30 d
- 0
Aún no hay ventas.
Ventas anónimas: sin comprador ni vendedor. Las cifras solo cuentan ventas entre jugadores.
En Wikipedia
Texto en inglés Aún no hay artículo en tu idioma: extracto en inglés.
In robot learning, a vision–language–action model (VLA) is a class of multimodal foundation models that integrates vision, language and actions. Given an input image (or video) of the robot's surroundings and a text instruction, a VLA directly outputs low-level robot actions that can be executed to accomplish the requested task. VLAs are generally constructed by fine-tuning a vision-language model (VLM) (i.e. a large language model extended with vision capabilities) on a large-scale dataset that pairs visual observation and language instructions with robot trajectories. These models combine a vision-language encoder (vision transformer), which translates an image observation and a natural language description into a distribution within a latent space, with an action decoder that transforms this representation into continuous output actions, directly executable on the robot. The concept was pioneered in July 2023 by Google DeepMind with RT-2, a VLM adapted for end-to-end manipulation tasks, capable of unifying perception, reasoning and control.
Texto: Wikipedia en inglés, CC BY-SA 4.0. ·
Cartas cercanas
Modelo–vista–modelo de vista
Patrón Modelo-Vista-Vista Modelo (MVVM)
Nº Q1247905 ★★
Aprendizaje supervisado
Tarea de aprendizaje automático de aprender una función que asigna una entrada a una salida basada en pares de entrada-salida de ejemplo
Nº Q334384 ★★
Sistema multiagente
Nº Q529909 ★★
Q-learning
Nº Q2664563 ★★★
DeepSeek-V4
Large language model
Nº Q139667467 ★★★★
Sycophancy (artificial intelligence)
Tendency of AI systems to tell users what they want to hear
Nº Q139915450 ★★