Date of Graduation

7-2026

Document Type

Thesis

Degree Name

Master of Science in Computer Science (MS)

Degree Level

Graduate

Department

Computer Science & Computer Engineering

Advisor/Mentor

Gauch, John

Committee Member

Zhang, Lu

Second Committee Member

Gauch, Susan

Keywords

3D Content, generation methods, benchmarking, unified evaluation framework

Abstract

In recent years, 3D generation has rapidly advanced with the development of powerful generative AI models capable of producing high-quality 3D content from various modalities, including text, images, and multi-view inputs. These advancements have significantly accelerated progress in applications such as gaming, virtual reality, robotics, and digital content creation. Despite this progress, there is still a lack of standardized and fair benchmarking protocols for evaluating 3D generation methods. Existing approaches are often assessed under inconsistent experimental settings, using different datasets, evaluation metrics, and processing pipelines. Such inconsistencies make reliable and objective comparisons difficult, limiting our understanding of the strengths and weaknesses of current methods and slowing the development of more robust and generalizable 3D generation systems. In this thesis, I analyze the limitations of existing benchmarking practices for 3D content generation and highlight the need for a unified evaluation framework. To address these challenges, I propose a large-scale text–image prompt dataset together with a consistent and fair evaluation protocol for benchmarking modern 3D generation methods. This work considers the two major paradigms in contemporary 3D generation research: optimization-based methods, which refine 3D representations through test-time optimization guided by a given prompt, and learning-based methods, which leverage large-scale datasets to train foundation models capable of efficient feed-forward 3D generation. Specifically, the proposed dataset consists of text–image pairs designed to reflect real-world objects and scenarios. The prompts are constructed from approximately 709,000 objects collected from real-world environments, improving the alignment between textual descriptions and realistic 3D structures. In the experiments, approximately 3,712 prompts are evaluated across six representative 3D generation methods. The experimental results demonstrate that the proposed dataset and benchmarking protocol provide more reliable, consistent, and interpretable evaluations of different approaches. In conclusion, this thesis provides the research community with a practical and standardized framework for benchmarking 3D generation methods, enabling more transparent comparisons and supporting future advancements toward more robust and generalizable 3D generation systems.

Share

COinS