Research Data and Analytical Materials for the Doctoral Dissertation: Health Behaviour, Preventive Healthcare Participation and Health Development Office Services Among Older Adults in Újbuda, Hungary
Description
Semmelweis University
Doctoral College
Health Sciences Division
Interdisciplinary Applied Health Sciences Program
Doctoral Dissertation Research Materials: Datasets, Statistical Analyses, and Supporting Files
This repository contains the research data and analytical materials associated with the doctoral dissertation entitled:
The Impact of the Újbuda Health Development Office on Health-Related Thinking, Values and Attitudes among Older Adults
(Az Újbudai Egészségfejlesztési Iroda működésének hatása az egészséggel kapcsolatos gondolkodásmódra, értékekre, attitűdökre az időskorú lakosság körében).
The study examines health behaviour, preventive attitudes, participation in health promotion programmes, and factors influencing the uptake of preventive services among adults aged 60 years and older in Újbuda (XI. district), Budapest, Hungary.
Questionnaires and study design
Surveys_2024_02_05.pdf – Collection of the original Hungarian-language questionnaires and the study-design overview used in the doctoral research. The document contains the Phase I PRE and POST questionnaires, the Phase II long-term retrospective follow-up questionnaire, and the Phase III cross-sectional questionnaire. The first page provides an overview of the study structure and the planned analytical relationships between the study phases.
In the study-design overview, “acceptors” refers to participants who had actually used the preventive services of the Újbuda Health Development Office. This group comprises the 500 participants in Phase I and the 100 former programme participants included in Phase II (total N = 600). The Phase III sample comprised adults aged 60 years or older recruited at the Újbudai Szent Kristóf Outpatient Clinic and its affiliated sites. It was used to examine awareness of the Health Development Office, willingness to participate in preventive programmes, and barriers to service use.
The questionnaires are provided in their original Hungarian-language form to document the measurement instruments and to support interpretation and reproducibility of the deposited datasets and statistical analyses.
The research consists of three phases:
I. Short-term prospective pre–post study:
A longitudinal questionnaire study among 500 adults aged 60 years or older who participated in a six-session health promotion programme of the Újbuda Health Development Office for the first time. The study assessed short-term changes in health-related attitudes, perceived importance of prevention, health behaviour, and well-being, as well as health behaviour profiles among new programme participants.
I_phase_PRE.sav – Anonymised SPSS dataset containing the baseline (PRE) questionnaire data of the 500 participants enrolled in Phase I of the study. The file contains the participants’ responses recorded before participation in the six-session health promotion programme and serves as the baseline dataset for descriptive analyses and participant profiling.
For the Phase I cluster analysis, the variables included in the clustering procedure were standardised as Z-scores prior to analysis in order to place variables measured on different numerical scales onto a comparable metric. This preprocessing step reduces the risk that variables with larger numerical ranges disproportionately influence the distance calculations used in the k-means clustering procedure.
For the income-related variable, a logarithmic transformation was additionally applied before standardisation because of its distributional characteristics. The resulting transformed variable (log_income_person) was subsequently standardised (Zlog_income_person) and used in the cluster-analysis dataset.
The dataset retains both the original and the derived standardised variables, supporting transparency and reproducibility of the preprocessing and subsequent cluster analyses.
I_phase_final_merged.sav – Anonymised merged SPSS dataset containing the matched PRE and POST questionnaire data of the 500 Phase I participants. The dataset was created for paired longitudinal analyses of changes following the health promotion programme. It contains the corresponding baseline and follow-up variables used, among others, for the Wilcoxon signed-rank analyses, effect-size estimation, and other PRE-POST comparisons.
Both datasets are anonymised and contain no directly identifying personal information.
I_Phase_Income_Score_Calculator_VBA.txt – Visual Basic for Applications (VBA) source code used to calculate the derived income-related score for participants in Phase I. The algorithm assigns predefined monetary values to selected questionnaire responses and sums these components into a single calculated value for each participant. The calculation uses information from multiple questionnaire variables, including categorical and numeric responses, and writes the resulting composite value into the designated output variable. The original SPSS dataset I_phase_PRE.sav was exported to Microsoft Excel solely as an intermediate processing step, where the VBA code was applied to calculate the derived income-related variables (@income and @income_person); the resulting variables were subsequently incorporated into the analytical dataset, and the per-person income-related variable was further log-transformed and standardised for the cluster analysis.
The score was calculated for Excel records 2–501, corresponding to the 500 participants included in Phase I (code was calculated in Excel environment). The VBA code is provided to document the exact scoring procedure and to support reproducibility of the derived variable used in subsequent statistical analyses.
Important: the calculated value should be interpreted as a constructed income-related score based on predefined scoring rules, rather than as a directly observed or verified individual income measure.
Statistical Results:
I_Phase_Wilcoxon_rank_test_and_biserial_bootstrap.zip
Compressed archive containing the complete analytical material for the Phase I H1 PRE-POST analysis. The archive includes the original SPSS output of the Wilcoxon signed-rank tests, a PDF export of the corresponding results, the SPSS syntax used to calculate paired-sample rank-biserial effect sizes, and 95% bootstrap confidence intervals based on 10,000 resamples, together with the resulting SPSS output and descriptive statistics for the analysed PRE and POST variables.
The analysis covers four health-related variables: perceived influence over one’s own health, importance of regular physical activity, importance of healthy nutrition, and importance of mental health and well-being. All analyses were based on paired PRE–POST observations from 500 participants in Phase I.
The Wilcoxon signed-rank tests were used to evaluate whether the PRE-POST differences were statistically significant, while paired-sample rank-biserial correlations were calculated to quantify effect size. Uncertainty around the effect-size estimates was assessed using 10,000 bootstrap resamples to obtain 95% confidence intervals.
The source dataset used for the Phase I Wilcoxon signed-rank, rank-biserial effect-size, and bootstrap analyses was I_phase_final_merged.sav.
Cluster_analysis_I_Phase.zip
Compressed archive containing the complete cluster-analysis output for Phase I. The archive includes separate SPSS output files for the 2-, 3-, 4-, and 5-cluster k-means solutions, together with the corresponding silhouette analyses used to evaluate cluster quality and support selection of the most appropriate cluster solution.
For each candidate solution, the archive contains the SPSS output of the k-means cluster analysis and the associated silhouette calculation. The analyses were performed to compare alternative cluster structures and to assess the internal cohesion and separation of the identified participant groups.
The archive also contains an Excel summary file (Silhouette_osszesites_I_Phase.xlsx) compiling the silhouette results across the 2-, 3-, 4-, and 5-cluster solutions to facilitate direct comparison of model performance.
The cluster analyses were based on the Phase I participant dataset and were used for the exploratory identification of heterogeneous participant profiles.
All k-means models were run with a maximum of 50 iterations to allow convergence of the cluster centres.
Although the two-cluster solution had the highest average silhouette coefficient, the three-cluster solution was retained for substantive interpretation because it provided more informative and interpretable profiles for the health-promotion target group.
These two SPSS output files document the descriptive statistical examination of the Phase I PRE-POST dataset and provide the underlying results used to characterise the study population and the distributions of the questionnaire variables before and after participation in the programme.
Pre_Descriptive.spv
SPSS Statistics output containing the descriptive analysis of the PRE measurements collected in Phase I. The output provides frequency distributions and summary statistics for the baseline questionnaire variables, including measures of central tendency and dispersion, quartiles, minimum and maximum values, and distributional characteristics such as skewness and kurtosis. Frequency-based graphical summaries are also included where applicable.
Post_Descriptive.spv
SPSS Statistics output containing the corresponding descriptive analysis of the POST measurements collected in Phase I. The same descriptive-statistical framework was applied to the post-intervention questionnaire variables to ensure comparability with the PRE measurements. The output includes frequency distributions, summary statistics, measures of central tendency and dispersion, quartiles, minimum and maximum values, skewness and kurtosis, together with frequency-based graphical summaries where applicable.
II. Long-term retrospective follow-up study:
A questionnaire-based follow-up study among 100 adults aged 60 years or older who completed at least six sessions of an Újbuda Health Development Office programme in 2019. The study examined health behaviour and preventive attitude patterns, together with self-reported longer-term changes among former programme participants.
II_Phase_final.sav
Anonymised SPSS dataset containing the questionnaire data of participants included in Phase II of the study. Phase II represents the long-term retrospective follow-up component and includes adults aged 60 years or older who had previously participated in at least six sessions of the Újbuda Health Development Office programme.
The dataset contains the variables used to examine health behaviour, prevention-related attitudes, and self-reported longer-term changes among former programme participants. The file serves as the source dataset for the descriptive analyses conducted in Phase II.
The dataset is anonymised and contains no directly identifying personal information.
Statistical Results:
II_Phase_Descriptive_stat.spv
SPSS Statistics output containing the descriptive statistical analysis of the Phase II dataset. The output provides frequency distributions and summary statistics for the questionnaire variables, including measures of central tendency and dispersion, quartiles, minimum and maximum values, and distributional characteristics such as skewness and kurtosis. Frequency-based graphical summaries are also included where applicable.
The descriptive analysis was used to characterise the Phase II study population and to examine the distributions and self-reported longer-term patterns of the questionnaire variables.
III. Cross-sectional study: A questionnaire survey of 500 adults aged 60 years or older recruited at the Újbudai Szent Kristóf Outpatient Clinic and its affiliated sites. This phase examined awareness of the Újbuda Health Development Office, willingness to participate in health promotion programmes, barriers to participation, and the relationship between the perceived importance of health maintenance and self-reported health-related activity.
The repository includes anonymised SPSS datasets, SPSS output files, and the SPSS syntax used for data management and statistical analyses. The materials are provided to support transparency, reproducibility, and the verification of the reported analyses.
III_Phase_final_2025_11_12b.sav
Anonymised SPSS dataset containing the questionnaire data of the 500 participants included in Phase III of the study. Phase III represents the cross-sectional component and includes adults aged 60 years or older recruited at the Újbudai Szent Kristóf Outpatient Clinic and its affiliated sites.
The dataset contains sociodemographic variables, self-reported health status, health-related behaviours, barriers to health maintenance, awareness of the Újbuda Health Development Office, willingness to participate in preventive programmes, and variables measuring the perceived importance of health maintenance and self-reported health-related activity.
For the Phase III cluster analysis, the variables included in the clustering procedure were standardised as Z-scores prior to analysis in order to place variables measured on different numerical scales onto a comparable metric. This preprocessing step reduces the risk that variables with larger numerical ranges disproportionately influence the distance calculations used in the k-means clustering procedure.
For the income-related variable, a logarithmic transformation was additionally applied before standardisation because of its distributional characteristics. The transformed income-related variable was subsequently standardised and used in the cluster-analysis dataset.
The dataset retains both the original variables and the derived transformed and standardised variables, supporting transparency and reproducibility of the preprocessing and subsequent cluster analyses.
The dataset is anonymised and contains no directly identifying personal information.
III_Phase_Income_Score_Calculator_VBA.txt
The original SPSS dataset III_Phase_final_2025_11_12b.savwas exported to Microsoft Excel solely as an intermediate processing step, where the VBA code was applied to calculate the derived income-related variables (@income and @income_person) for the 500 participants included in Phase III. The resulting variables were subsequently incorporated into the analytical SPSS dataset. The per-person income-related variable was then logarithmically transformed (log_person_jovedelm) because of its distributional characteristics and subsequently standardised for use in the cluster analysis.
The VBA code is provided to document the exact scoring procedure and to support reproducibility of the derived variables used in subsequent statistical analyses.
Important: the calculated values should be interpreted as constructed income-related measures based on predefined scoring rules, rather than as directly observed or verified individual income measures.
Statistical Results:
Descriptive_.spv
SPSS Statistics output containing the descriptive statistical analysis of the Phase III dataset. The output provides frequency distributions, summary statistics, and graphical summaries for the questionnaire variables used in the cross-sectional study.
The analyses cover sociodemographic characteristics, self-reported health status and chronic health conditions, mobility, barriers to health maintenance, environmental and transport-related factors, internet access, awareness of the Újbuda Health Development Office, willingness to participate in preventive programmes, and interest in specific health promotion services. The output also includes descriptive information on the perceived importance of health maintenance, self-reported health-related activity, anthropometric variables, and the derived income-related measures.
Frequency-based bar charts are included for the analysed variables where applicable.
The descriptive analysis was used to characterise the Phase III study population and to examine the distributions of the questionnaire variables prior to subsequent comparative and cluster-based statistical analyses.
III_Phase_Cluster_Analysis_3_Clusters.spv
SPSS Statistics output containing the cluster analysis conducted in Phase III of the study.
The analysis was used to identify distinct participant profiles based on standardised sociodemographic, socioeconomic, health-related, environmental, service-use, and health-behaviour variables. The variables included in the clustering procedure were transformed to Z-scores prior to analysis in order to place them on a comparable numerical scale. The logarithmically transformed indirect socioeconomic status indicator was also standardised before inclusion in the cluster model.
Hierarchical cluster analysis was first used to support the assessment of the appropriate cluster structure, followed by K-means clustering to obtain the final participant classification. The uploaded output documents the selected three-cluster solution.
The final K-means model was based on participants with complete data for all variables included in the clustering procedure. The output contains the initial and final cluster centres, iteration history, cluster membership, distances between final cluster centres, descriptive between-cluster comparisons, and the number of cases assigned to each cluster.
The final three-cluster solution was subsequently interpreted as distinct health-behavioural and socioeconomic profiles within the Phase III study population. Internal validation and comparison of alternative 2-, 3-, 4-, and 5-cluster solutions using silhouette analysis were performed separately and are documented in the corresponding repository files.
III_Phase_Cluster_analyses.zip
Compressed archive containing the supplementary analytical materials used to assess and validate the number of clusters in the Phase III cluster analysis.
The archive includes the alternative 2-, 3-, 4-, and 5-cluster solutions and the corresponding silhouette-analysis outputs used to compare their internal cluster separation.
The alternative solutions were evaluated using average silhouette coefficients based on Euclidean distance. The three-cluster solution showed the highest average silhouette value among the evaluated models and was therefore retained as the final solution, together with considerations of cluster size and substantive interpretability.
These files are provided as supplementary documentation of the cluster-selection and internal-validation procedure. The detailed output of the final three-cluster K-means model is provided separately in III_Phase_Cluster_Analysis_3_Clusters.spv.
III_Phase_Wilcoxon_Rank_Biserial_Bootstrap.zip
Compressed archive containing the complete analytical materials for the Phase III comparison between the perceived importance of maintaining health and self-reported health-maintenance activity.
The archive includes the original SPSS output of the Wilcoxon signed-rank test, the SPSS syntax used to calculate the paired-sample rank-biserial correlation and its 95% bootstrap confidence interval, and the corresponding SPSS output containing the effect-size results.
The analysis was based on paired observations from the 500 participants included in Phase III. The Wilcoxon signed-rank test was used to evaluate whether a systematic difference existed between the perceived importance of health maintenance and the level of self-reported activity undertaken to maintain health.
The magnitude of this difference was quantified using paired-sample rank-biserial correlation. Uncertainty around the effect-size estimate was assessed using 10,000 bootstrap resamples, with a fixed random seed to support computational reproducibility.
The archive is provided to document both the statistical significance testing and the corresponding effect-size estimation, and to support transparent reproduction of the Phase III value–action discrepancy analysis.
The source dataset used for these analyses was III_Phase_final_2025_11_12b.sav.
Software: Statistical analyses were conducted using IBM SPSS Statistics version 30.0.0.0 (171). Microsoft Excel was used for selected data-processing procedures.
Participant_Information_and_Data_Protection_Notice_Phases_I-II-III.pdf
Participant information and data protection notice covering all three phases of the study. The document provides participants with information on the purpose and conduct of the research, voluntary participation, data handling and protection, confidentiality, and the use of collected data for scientific research purposes.
The notice also describes the study-specific data handling procedures applicable to Phases I, II, and III and provides information on the ethical framework of the research. The study was conducted in accordance with the principles of the Declaration of Helsinki and was reviewed and approved by the Regional and Institutional Research Ethics Committee of the South Buda Central Hospital – Szent Imre University Teaching Hospital (protocol/minutes no. 9/2024; 2 May 2024).
The document is provided as part of the repository materials to document the participant information and data protection procedures applied in the doctoral research and to support transparency and reproducibility of the study documentation.
Ethical approval: The study was reviewed by the Regional and Institutional Research Ethics Committee of the South Buda Central Hospital – Szent Imre University Teaching Hospital (protocol/minutes no. 9/2024; 2 May 2024).
Contact: For questions regarding the dataset or analytical materials, please contact Peter Domjan at domjan.peter@phd.semmelweis.hu or peterdomjan@gmail.com
Files
Cluster_analysis_I_Phase.zip
Files
(3.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:e8d15fbfac43361857ef96479f3fc7b8
|
203.0 kB | Preview Download |
|
md5:b7617af7aa8149fa2bee7fc72e9a4a61
|
246.5 kB | Download |
|
md5:5782e6ac67081de85d54df7d81221d79
|
100.3 kB | Download |
|
md5:14f870f5ef398541f377bf668185be0c
|
4.4 kB | Preview Download |
|
md5:f18704f430586667953bf4cf71148bb2
|
168.7 kB | Download |
|
md5:bbeaf510f6f6bf4e2590457e7ac918fa
|
144.5 kB | Preview Download |
|
md5:0a043be4d38b092adffe963b7c476c5f
|
129.0 kB | Download |
|
md5:13ba797193130ac533eab85f303a1fe5
|
11.2 kB | Download |
|
md5:f1cc1e813ee06ffba24f2ed56ee05385
|
1.0 MB | Preview Download |
|
md5:50b9cfdafa4ac52db049c19d275c242d
|
73.2 kB | Download |
|
md5:54ddae6ba4a74653c96ada02ddcfb19a
|
114.8 kB | Download |
|
md5:feb74a97c67bdd8bcc05c75bab742404
|
6.1 kB | Preview Download |
|
md5:9b33da43137ea298f55b5930f6250963
|
8.7 kB | Preview Download |
|
md5:e09be320ede9f9f046a42ea447af0519
|
136.5 kB | Preview Download |
|
md5:58668e28d8d1db818022485c52d727b8
|
79.7 kB | Download |
|
md5:709c79f88738c840e5a95c1b5fa7a484
|
384.4 kB | Download |
|
md5:09f17c8d9243b79b9f732d64b47fb151
|
22.0 kB | Preview Download |
|
md5:607586728ae0e2634921559e133aac67
|
345.6 kB | Preview Download |
Additional details
Funding
- Ministry for Culture and Innovation
- University Research Scholarship Programme – Cooperative Doctoral Programme 2024 (EKÖP-KDP-2024) 2024-2.1.2-EKÖP-KDP-2024-00002
Dates
- Collected
-
2025-06-01/2025-12-31Phase I – short-term prospective PRE–POST data collection
- Collected
-
2025-01-01/2025-12-31Phase II – long-term retrospective follow-up data collection
- Collected
-
2025-01-01/2025-12-31Phase III – cross-sectional data collection