Survey of the Global Fortran Community
Authors/Creators
Description
This dataset contains the survey responses and associated analysis artefacts collected for the study “What do we really know about Fortran?”. The study investigates the current state of Fortran codebases in scientific and engineering computing, with a particular focus on the long-term sustainability challenges faced by the practitioners who maintain them.
The survey was conducted in 2025–2026 and gathered responses from 150 participants across 25 countries. Respondents were asked to focus on a single Fortran codebase of their choosing, and to answer 39 questions spanning 8 thematic sections: codebase characteristics (size, age, Fortran standard versions, licensing), technical dependencies and HPC requirements, documentation practices, legacy status, key challenges, potential improvements, organisational context, scientific domain, and respondent demographics.
The codebases described by respondents support work in a wide range of safety- and mission-critical domains, including numerical weather prediction, nuclear physics, plasma physics and fusion energy research, geophysics, computational chemistry, atmospheric science, and national defence. Codebase sizes range from under 1,000 lines of code to more than 5 million lines. The majority of respondents (82%) hold a PhD or higher qualification, 32% are over 55 years of age, and 26% are the sole maintainer of their codebase — together these figures underscore the workforce sustainability risks that motivated the study.
The dataset is provided to support the reproducibility of the analyses reported in the manuscript and to enable further research on the long-term sustainability of scientific software. High resolution images from the published paper are also included for convenience.
Methods
Survey
The survey was administered online. Participants were recruited through the Fortran-lang community, the BCS Fortran Specialist Group, social media, and direct contact with researchers known to work with Fortran codebases. Participation was voluntary and anonymous; informed consent was obtained from each respondent before any data were collected.
The survey comprised 39 questions across 8 sections:
- Codebase selection and scope
- Codebase characteristics (size, language composition, HPC requirements)
- Dependencies and licensing
- Documentation
- Legacy status and criteria
- Challenges and potential improvements (free-text)
- Organisational and disciplinary context
- Respondent demographics
The full survey is provided as a separate PDF file in this dataset.
Data collection and cleaning
Raw survey responses were exported directly from the survey platform and are preserved without alteration in the RAW survey data sheet. Four free-text or semi-structured columns — codebase size (lines of code), respondent role, organisation type, and highest academic qualification — required manual or semi-automated parsing to produce usable categorical or numeric values; these cleaned values are provided in the CLEANED categories sheet alongside the original free-text entries for transparency. The free-text responses to the challenge question were independently coded by two researchers using a set of 15 thematic categories; the coded responses are stored in the Challenges Categories sheet.
Table of contents
Files in This Dataset
|
File |
Description |
|
Fortran_survey_data.xlsx |
Primary data file containing all survey responses and derived analyses (see sheet descriptions below) |
|
Fortran Survey Form as PDF (redacted).pdf |
The survey as presented to participants, showing question wording, order, and response options |
|
FortranFuture Consent Form (redacted).pdf |
The consent agreement approved by the institutions ethical approval board |
|
FortranFuture Survey Participant Information Sheet v3 (redacted).pdf |
Infomation provided to survey participants |
|
FortranFuture Participant README (redacted).pdf |
Invitation to participate in the survey |
|
survey_analysis.py |
python code to produce bubble plots/histograms included in the paper |
|
plot_codebase_sizes.py |
python code to produce the box plot of codebase sizes included in the paper |
|
age_by_gender.pdf |
Histogram showing age of respondents broken down by gender |
|
bus_factor.pdf |
Bubble plot showing the respondent's age vs the number of people currently working on the codebase |
|
codebase_sizes.pdf |
Box plot of the numerical size (LOC) vs the ordinal size of the respondents' codebases |
|
codebase-size-vs-team-size.pdf |
Bubble plot showing numerical size of respondents' codebases vs the number of people currently working on the codebase |
|
framework.pdf |
A framework for studying routes to impact of Fortran |
|
qualification.pdf |
Bubble plot showing the respondents' highest qualification vs the relevance of that qualification to their work |
|
regions.pdf |
Histogram showing the geographical location of the respondents |
|
role_vs_owner.pdf |
Type of organisation that owns the codebase vs the role of the respondent in that organisation |
Data File Structure: Fortran_survey_data.xlsx
The workbook contains three sheets.
Sheet 1 — RAW survey data
150 rows × 15 columns. Contains one row per respondent with the original survey responses, exported directly from the survey platform. Only the columns used in the analyses reported in the manuscript are included.
|
Column |
Note |
|
ID |
Unique integer identifier for each response (used to cross-reference sheets) |
|
How many Fortran codebases do you work on... |
Multiple-choice selection (Q2) |
|
Approximately, how many people currently work on the codebase... |
Multiple-choice selection (Q6) |
|
What is the approximate size of the current codebase... |
Free-text entry (Q7) |
|
In subjective terms, how would you characterise the size of the codebase? |
Multiple-choice selection (Q8) |
|
Do you personally consider the codebase to be a legacy codebase? |
Multiple-choice selection (Q21) |
|
Do others you work with consider the codebase to be a legacy codebase? |
Multiple-choice selection (Q22) |
|
What is the most significant challenge...you personally confront with the codebase? |
Free-text entry (Q24) |
|
In what country is the organisation you work for located? |
Multiple-choice selection (Q26) |
|
What type of organisation "owns" the codebase? |
Multiple-choice selection (Q27) |
|
What is your main role in the organisation... |
Multiple-choice selection (Q28) |
|
What is your age? |
Multiple-choice selection (Q32) |
|
What is your gender? |
Multiple-choice selection (Q33) |
|
What is your highest academic qualification? |
Multiple-choice selection (Q34) |
|
How relevant is your highest qualification to the work you do on, or with, this codebase? |
Multiple-choice selection (Q35) |
Sheet 2 — CLEANED categories
150 rows × 8 columns. This sheet provides cleaned, machine-readable versions of four free-text or ambiguous columns from the raw data, alongside the original text for reference.
|
Column |
Content |
|
ID |
Unique integer identifier for each response (used to cross-reference sheets) |
|
What is the approximate size of the current codebase... |
Original free-text response (copied from raw) |
|
CLEANED Codebase LOC |
Numeric Lines of Code value parsed from the free-text response (Blank where unparseable) |
|
What is your main role in the organisation... |
Original free-text role response (copied from raw) |
|
CLEANED Role |
Standardised role category |
|
What type of organisation "owns" the codebase? |
Original organisation response (copied from raw) |
|
CLEANED Organisation |
Standardised organisation type category |
|
What is your highest academic qualification? |
Original free-text qualification response (copied from raw) |
|
CLEANED Qualification |
Standardised qualification category |
Sheet 3 — Challenges Categories
153 rows × 18 columns. Contains the coded results of a thematic analysis applied to free-text responses to the challenge question (Q24).
|
Column |
Content |
|
ID |
Unique integer identifier for each response (used to cross-reference sheets) |
|
Response |
The verbatim free-text response that was coded |
|
Complexity |
Binary flag |
|
Documentation |
Binary flag |
|
Legacy & Technical Debt |
Binary flag |
|
Testing & Debugging |
Binary flag |
|
Compiler & Language support |
Binary flag |
|
GPU & HPC |
Binary flag |
|
Human Resources |
Binary flag |
|
Time & Funding |
Binary flag |
|
Perception |
Binary flag |
|
Portability |
Binary flag |
|
Interoperability with other codebases & languages |
Binary flag |
|
Dependency management & build systems |
Binary flag |
|
Domain knowledge |
Binary flag |
|
Organisational factors |
Binary flag |
|
Software Engineering practices & tooling |
Binary flag |
|
Null response |
Binary flag |
A single response may be flagged under multiple thematic categories.
Technical info
Technical Details
File format: Microsoft Excel Open XML Format (.xlsx). The file can be opened with Microsoft Excel, LibreOffice Calc, Google Sheets, or read programmatically using libraries such as pandas (Python), readxl (R), or equivalent tools in other environments.
Character encoding: The .xlsx file uses UTF-8 encoding throughout. Note that a small number of column header values in the raw data contain non-breaking space characters (\xa0) as exported by the survey platform; these are preserved as-is.
Missing values: Missing values are represented as blank cells. The CLEANED Codebase LOC column contains Blanks for responses that could not be parsed to a numeric value (e.g., free-text descriptions without a number).
Files
age_by_gender.pdf
Files
(1.3 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:dd38a83e2b4a97fb52ed345af4dc0f28
|
20.1 kB | Preview Download |
|
md5:ea1997cbeebcef3d11215e6cdc0dc7de
|
23.9 kB | Preview Download |
|
md5:305b515ca7d1c3310f1633a70dc640bc
|
24.4 kB | Preview Download |
|
md5:4ab9560b6ec12ea4b16a80bcf99b9812
|
27.7 kB | Preview Download |
|
md5:6d467f8f066c1daf9be6393f54a04e52
|
298.9 kB | Preview Download |
|
md5:4c19be6687f86de79971a4f88d8d7de8
|
62.7 kB | Download |
|
md5:fc10406509e50980f0a2464fbf12342a
|
254.9 kB | Preview Download |
|
md5:d6376368600bb6180e6c0b1f107ac255
|
39.4 kB | Preview Download |
|
md5:e11d342963c79c6762c98f4e25c7545f
|
295.1 kB | Preview Download |
|
md5:05a44343e16b5999fca263bb71eb1e94
|
15.0 kB | Preview Download |
|
md5:7df10b49b0c9191b6707dd4d7c0ab005
|
6.2 kB | Download |
|
md5:93cb854faf4f7d857bdbbcb74079ebee
|
22.1 kB | Preview Download |
|
md5:fd7208be32c3cc16d80455c2111cbe62
|
127.9 kB | Preview Download |
|
md5:35594be04c89c9db5cbd57ec6ef7aa34
|
26.2 kB | Preview Download |
|
md5:5813780412ee42c3b4a462bd436df788
|
33.8 kB | Download |