Monitor this situation.
Unsubscribe anytime.
[SITUATION] · [ACTIVE] · [TECHNOLOGY]
2 clusters · 4 sources · 1 days · First seen · Last updated
AI efficiency and latency reduction research
Overview
Researchers are developing new methods to improve the efficiency and speed of artificial intelligence systems.
In robotics, a collaborative effort involving MIT and Nvidia introduced VLASH (Vision-Language-Action with Scheduled Heuristics). This method allows robots to calculate subsequent commands while performing current tasks, reducing response latency by up to 17.4 times and doubling the speed of specific tasks like cube sorting.
In the field of large language models, Nvidia researchers introduced a cross-model KV cache transfer technique. This method aims to reduce the computational costs and latency in multi-model workflows by mapping prefilled data from a source model to a target model, rather than recomputing conversation history. This process can run between 2.7 and 25 times faster than recomputation while maintaining approximately 98% of a model’s standalone accuracy.
Entities
Timeline
-
5 days ago
[TECHNOLOGY] 2 sourcesNvidia researchers develop method to reduce AI model handoff costsNvidia researchers developed a linear math technique for cross-model KV cache transfer, significantly reducing latency and costs in multi-LLM workflows while maintaining high accuracy.
-
6 days ago
[TECHNOLOGY] 2 sourcesMIT and Nvidia researchers develop AI to reduce robot latencyScientists from MIT, Nvidia, and Caltech developed VLASH, an AI method that reduces robotic latency by up to 17.4 times by allowing robots to calculate next steps while performing current tasks.
Sources
cafebiz.vn · cafef.vn · dev.to · venturebeat.com