Abstract
Large-scale image and video datasets are indispensable for the development of computer vision algorithms,mainly because they provide the necessary resources to train and evaluate various models. Constructing such datasets for different computer vision tasks is a crucial but complex task,because it involves considerable challenges in data collec-tion,annotation,and preservation of data diversity. Traditionally,acquiring large,high-quality image and video datasets has been a resource-intensive task,requiring manual labeling,data collection in real-world settings,and the use of special-ized hardware for capturing high-quality images and videos. As deep learning methods increasingly rely on large-scale labeled data,the need for innovative data generation techniques has become more prominent. In recent years,generative models,such as generative adversarial network(GAN)and diffusion models have emerged as powerful tools for generating synthetic datasets. These models can create diverse,controllable,and highly realistic image and video data,offering an effective alternative or supplement to traditional data collection methods. By using these techniques,vast amounts of data can be generated to represent various scenarios and conditions,which are essential for training robust computer vision mod-els. Generative models provide a flexible solution that can generate data without the need for real-world data acquisition,unlike traditional data collection,which is often constrained by geographic,financial,and logistical limitations. This review begins by introducing the significance and background of image and video data generation in computer vision. Image and video data play a critical role in the development and training of computer vision algorithms,as large-scale,diverse datasets are essential for building robust models. Moving on,the review categorizes the key data generation techniques into three broad approaches:traditional data augmentation methods,3D rendering-based generation methods,and deep genera-tive models. First,traditional data augmentation techniques,including geometric transformations,color adjustments,and cropping,are commonly used to improve model generalization by expanding existing datasets. Although these methods are relatively simple and computationally inexpensive,their ability to generate diverse and realistic datasets is limited. In com-parison,3D rendering technologies,such as virtual engines and neural radiance fields(NeRF),enable the creation of highly realistic synthetic data by simulating real-world environments. These technologies have the advantage of generating diverse datasets by adjusting environmental factors,such as lighting,object interactions,and camera angles. Further-more,deep generative models,such as GAN and diffusion models,have shown remarkable effectiveness in generating high-quality synthetic data. On the one hand,GAN work by training two neural networks in a competitive manner in which a generator creates synthetic data,while a discriminator evaluates its realism. Over time,as the generator improves its out-put,it creates increasingly realistic data. Diffusion models,on the other hand,iteratively refine noisy data into clear and realistic images or videos,enabling the generation of diverse,high-quality datasets. Next,the review discusses the diverse applications of these generative models across a wide range of computer vision tasks. These tasks include image enhance-ment,object detection,tracking,pose or action recognition,biometric identification,crowd behavior analysis,and more recently,emerging fields like autonomous driving and embodied artificial intelligence. In particular,synthetic data have been instrumental in training models for tasks that are challenging to address using only real-world data. For example,in biometric identification,synthetic data can generate a wide variety of samples for fingerprints,faces,irises,and palm-prints,thus providing more diverse training examples and reducing reliance on real biometric data,which are often difficult to acquire. Similarly,in autonomous driving,synthetic data can generate various driving scenarios,including different road conditions,weather patterns,and traffic behaviors,thereby training autonomous vehicle models in safe and controlled environments. In recent years,synthetic data have also been proven invaluable in fields like pose and action recognition,where diverse datasets are essential for accurately detecting human actions across different settings and contexts. However,despite the considerable progress made in image and video data generation,several challenges remain. One of the primary issues is ensuring the realism and diversity of generated data,which is crucial for training models that can generalize well to real-world scenarios. Furthermore,despite significant advances in generative models,there remains a lack of research on how to effectively evaluate the quality of synthetic data and use feedback mechanisms to guide the generation process. In addition,ethical considerations surrounding the use of synthetic data,especially in sensitive applications,such as biomet-ric recognition,must be carefully addressed. Among others,the use of synthetic data raises concerns regarding privacy,consent,and potential misuse,which must be handled responsibly. Looking ahead,as generative models continue to evolve,they are expected to produce even more realistic and diverse datasets,thus offering new possibilities for training computer vision models. The future of image and video data generation holds great promise,with advancements in genera-tive technologies poised to drive further innovation in computer vision,AI,and many other fields.
| Translated title of the contribution | 面向计算机视觉的数据生成与应用研究进展 |
|---|---|
| Original language | English |
| Pages (from-to) | 1872-1952 |
| Number of pages | 81 |
| Journal | Journal of Image and Graphics |
| Volume | 30 |
| Issue number | 6 |
| DOIs | |
| State | Published - Jun 2025 |
Keywords
- 3D rendering
- autonomous driving
- biometric recognition
- computer vision
- conventional data generation
- crowd analysis
- data generation and application
- deep genera-tive model
- embodied artificial intelligence
- image enhancement
- individual analysis
- video generation
Fingerprint
Dive into the research topics of 'Recent advances in data generation and its applications in computer vision'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver