Abstract: Near-miss traffic incidents, positioned just above "unsafe acts" on the safety triangle theory, offer crucial predictive insights for preventing crashes. However, these incidents are often ...
Abstract: Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and ...
SlowFast-LLaVA is a training-free multimodal large language model (LLM) for video understanding and reasoning. Without requiring fine-tuning on any data, it achieves comparable or even better ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results