吴 飞,刘亚楠,庄越挺.基于张量表示的直推式多模态视频语义概念检测.软件学报,2008,19(11):2853-2868 |
基于张量表示的直推式多模态视频语义概念检测 |
Transductive Multi-Modality Video Semantic Concept Detection with Tensor Representation |
投稿时间:2008-03-01 修订日期:2008-08-26 |
DOI: |
中文关键词: 多模态 张量镜头 时序关联共生 高阶SVD 降维 直推式支持张量机 |
英文关键词:multi-modality TensorShot temporal associated cooccurrence (TAC) higher order SVD (HOSVD) dimensionality reduction transductive support tensor machine (TSTM) |
基金项目:Supported by the National Natural Science Foundation of China under Grant Nos.60603096, 60533090 (国家自然科学基金); the National High-Tech Research and Development Plan of China under Grant No.2006AA010107 (国家高技术研究发展计划(863); the National Key Technology R&D Program of China under Grant No.2007BAH11B01 (国家科技支撑计划); the Program for Changjiang Scholars and Innovative Research Team in University of China under Grant Nos.IRT0652, PCSIRT (长江学者和创新团队发展计划) |
|
摘要点击次数: 6792 |
全文下载次数: 7604 |
中文摘要: |
提出了一种基于高阶张量表示的视频语义分析与理解框架.在此框架中,视频镜头首先被表示成由视频中所包含的文本、视觉和听觉等多模态数据构成的三阶张量;其次,基于此三阶张量表达及视频的时序关联共生特性设计了一种子空间嵌入降维方法,称为张量镜头;由于直推式学习从已知样本出发能对特定的未知样本进行学习和识别,最后在这个框架中提出了一种基于张量镜头的直推式支持张量机算法,它不仅保持了张量镜头所在的流形空间的本征结构,而且能够将训练集合外数据直接映射到流形子空间,同时充分利用未标记样本改善分类器的学习性能.实验结果表明,该方法能够有效地进行视频镜头的语义概念检测. |
英文摘要: |
A higher-order tensor framework for video analysis and understanding is proposed in this paper. In this framework, image frame, audio and text are represented, which are the three modalities in video shots as data points by the 3rd-order tensor. Then a subspace embedding and dimension reduction method is proposed, which explicitly considers the manifold structure of the tensor space from temporal-sequenced associated co-occurring multimodal media data in video. It is called TensorShot approach. Transductive learning uses a large amount of unlabeled data together with the labeled data to build better classifiers. A transductive support tensor machines algorithm is proposed to train effective classifier. This algorithm preserves the intrinsic structure of the submanifold where tensorshots are sampled, and is also able to map out-of-sample data points directly. Moreover, the utilization of unlabeled data improves classification ability. Experimental results show that this method improves the performance of video semantic concept detection. |
HTML 下载PDF全文 查看/发表评论 下载PDF阅读器 |