| This study examines GPT model output reproducibility in qualitative data analysis by measuring output consistency across trials, model variants, and data types. Three GPT-5 variants (GPT-5, GPT-5-Thinking, and GPT-5-Pro) were applied to four publicly available qualitative data samples (legislation, article, interview transcript, and stakeholder assessment) under two prompt conditions: an original unperturbed prompt and a perturbed paraphrased version. Reproducibility was assessed using a multi-dimensional framework utilizing cosine similarity and Euclidean distance of output embeddings to measure semantic stability, and pairwise word-count difference to measure structural variability. Five independent outputs were generated per cell and reduced to a single median stability estimate per data type (n = 4) to support nonparametric inference. Prompt sensitivity was assessed using Wilcoxon signedrank tests, between-model mean rank differences using Friedman tests, and the relationship between stability dimensions using an exploratory Spearman correlation. Cosine similarity was consistently high across all models and conditions, corroborated by Euclidean distances, indicating stable semantic outputs under repeated runs. Word-count difference showed greater variability, highlighting a distinction between semantic and structural reproducibility. The Wilcoxon and Friedman tests did not reach statistical significance, reflecting limited power. While exploratory, the Spearman correlation was significant (n = 12 for individual conditions; n = 24 combined), supporting retention of both metrics and suggesting nonsignificant results reflect power limitations from a limited sample size rather than an absence of effects. Descriptive patterns suggest condition-dependent stability differences across model variants. The study contributes a multi-metric framework for evaluating LLM reproducibility in qualitative research contexts. Keywords: LLM reproducibility, LLM-assisted qualitative data analysis, GPT output stability, Prompt sensitivity, Semantic similarity, Structural variance |