byted-las-long-video-understand
Performs deep AI-powered analysis and understanding of long-form videos (up to 3 hours, 10GB) using Volcengine LAS large language models. Video analysis, video comprehension, and video summarization — generates comprehensive video summaries, video recaps, chapter breakdowns, event timelines, key moments, and structured content indexing. Supports behavior detection, action detection, video annotation, video tagging, and video content recognition. Enables intelligent video question answering — ask questions about video content and get AI answers. Handles long meeting recordings, lecture videos, webinars, surveillance and security footage, movies, tutorials, and any long video that needs detailed understanding. Async processing with submit-poll workflow. Use this skill when the user wants to analyze or understand long videos (up to 3h/10GB) with LLM-based deep comprehension, summarize video content or generate recaps, extract chapters/key moments/timelines from videos, do video Q&A (ask questions about video content), detect actions/behaviors in videos, review meeting recordings/lectures/webinars, or get structured video descriptions.