Skip to main content
v2026.11,858 entries · CC-BY 4.0

What Is Software Engineering? Research Software, Reproducibility, and Citation

A thorough answer to “what is software engineering”: its subfields, methods and history, plus research software engineers, reproducibility, software citation and who funds the work.

Written and maintained by CASRAI Editorial Board

Last updated

Software engineering is the discipline of applying systematic, measurable and disciplined methods to the specification, design, construction, testing, deployment and maintenance of software. Where programming is the act of writing code, software engineering is the broader practice of making software that is correct enough, maintainable enough and affordable enough to be relied on by people other than its author, and of keeping it that way as requirements, platforms and teams change. The field is codified by the IEEE Computer Society in the Guide to the Software Engineering Body of Knowledge (SWEBOK); version 4.0, published in 2024, organizes the discipline into 18 knowledge areas and added three new ones: software architecture, software engineering operations and software security.

For a research-administration audience the discipline matters for a specific reason: most research now runs on software. Analysis pipelines, simulation codes, instrument-control systems, statistical scripts and data-processing workflows all sit between raw observations and published results, and the quality of that software shapes whether a result can be checked, reproduced and reused. This guide explains what software engineering is, how it differs from adjacent fields, what its subfields and methods are, who funds the research, and how the emerging profession of the research software engineer (RSE) connects the discipline to the practical problems of reproducibility and software citation.

What software engineering studies

Software engineering asks how to build software reliably and at scale. Its recurring questions are practical ones. How do we establish what the software is supposed to do, and check that it does it? How should a large system be divided into parts that can be understood and changed independently? How do we find defects before users do, and how do we know when testing is sufficient? How do we change a system that is already in use without breaking it? How do teams of people coordinate on a shared codebase over years?

The discipline is empirical as well as practical. A large part of software engineering research studies real projects, such as defect data, code review practices, developer productivity, build failures and the effect of tooling on outcomes, using methods borrowed from statistics, the social sciences and experimental design. It is therefore a field where evidence about what works is actively contested and continually updated, rather than a settled body of rules.

A short history

The term “software engineering” is most closely associated with a NATO-sponsored conference held in Garmisch, Germany, in October 1968, which brought together more than fifty computer scientists and practitioners from eleven countries. The conference title was chosen deliberately to provoke the idea that software production needed the discipline associated with engineering, in response to widespread concern about projects running late, over budget and unreliable. The phrase is generally attributed to the conference chair, Friedrich L. Bauer, though historians of computing note that the term was in use before the conference and that the common claim it was coined there is an oversimplification.

From the 1970s onward the field developed structured programming, formal specification, software process models, and later object-oriented design, agile methods and DevOps. Version control, automated testing, continuous integration and code review became standard practice, and are the same techniques that research software engineers now advocate for scientific code. Professional codification followed through IEEE and ACM curricula and the SWEBOK. Accreditation of undergraduate software engineering programs is handled in the US by ABET, which has program criteria specific to software engineering.

Major subfields

  • Requirements engineering — eliciting, analyzing and documenting what a system must do, including the non-functional qualities such as performance, safety and usability.
  • Software design and architecture — decomposing a system into components and interfaces, and reasoning about the trade-offs that follow.
  • Software testing and verification — unit, integration and system testing, property-based and fuzz testing, static analysis, and formal verification for the most safety-critical code.
  • Software maintenance and evolution — refactoring, managing technical debt, migrating legacy systems, and understanding why software ages.
  • Software process and project management — agile, plan-driven and hybrid approaches, estimation, risk management and team organization.
  • Configuration management and operations — version control, build systems, release engineering, continuous integration and deployment, and monitoring in production.
  • Software security — secure development practices, vulnerability management and supply-chain integrity.
  • Empirical software engineering and mining software repositories — the study of how software is actually developed, using project data and controlled studies.
  • Human and social aspects — developer experience, collaboration, onboarding, ethics and the effects of tools on people.
  • Software engineering for specific domains — embedded and safety-critical systems, scientific computing, machine-learning systems, and regulated software such as medical-device software.

How software engineering differs from neighboring fields

Computer science studies computation itself: algorithms, complexity, languages, systems and the theory behind them. Software engineering applies that knowledge to the practical construction and long-term upkeep of software. The two overlap heavily and are often taught in the same departments; see the CASRAI guide to what computer science is for the theoretical side.

Systems engineering manages the design of complex systems as a whole, including hardware, software, people and processes. Software engineering is typically one component of that effort; see what systems engineering is. Electrical and mechanical engineering meet software engineering in embedded systems, controls and robotics, where code is tightly coupled to physical hardware; see electrical engineering and mechanical engineering. Data science extracts conclusions from data and leans on software engineering practices to make its pipelines dependable; see what data science is. For the parent category, see the guide to what engineering is.

Software engineering also differs from traditional engineering in its raw material. Software has no physical wear, can be copied at no cost and can be changed after delivery, which makes iteration cheap but also makes the accumulation of unmanaged change a defining difficulty of the field.

Software in science

Research increasingly depends on code that the researchers themselves wrote. A typical project may combine community libraries, a custom analysis script, a workflow manager and a computing cluster, and the published result is only as trustworthy as that stack. Several features distinguish scientific software from commercial software. It is often written by domain scientists whose training is in their discipline, not in software practice. Its requirements change as the research question changes. Its users are frequently a small group, sometimes only the author. And its output is knowledge, so a silent defect can produce a plausible but wrong scientific conclusion without any visible failure.

Software that underpins research ranges from short analysis scripts to long-lived community infrastructure maintained by many contributors. The engineering effort appropriate to each is different, which is one reason the research community distinguishes among throwaway scripts, shared lab tools and sustained community software, and plans maintenance accordingly. The open source software model is the dominant way community research infrastructure is developed and shared, typically hosted in a code repository that records the full history of changes.

Research software engineering and the RSE career

A research software engineer combines software expertise with an understanding of the research process. The term was coined in 2012 in the UK, following a workshop run by the Software Sustainability Institute at the University of Oxford that discussed the lack of career paths for people who write software in academia. The first dedicated RSE group was established at University College London in 2012, a UK association of research software engineers formed in 2013 and later became the Society of Research Software Engineering, and the first RSE conference took place in the UK in 2016. The US Research Software Engineer Association (US-RSE) grew out of a grassroots effort that began in 2017-2018 and held its first national conference in 2023. Comparable communities exist in Germany (de-RSE) and in many other countries.

RSEs fill a structural gap. Postdoctoral researchers and graduate students write much of the code in academic projects, but those positions are temporary and rewarded for publications, not for software quality or longevity. RSEs are typically employed in permanent or long-term roles, often in central research-computing units or embedded in project teams, where they build and maintain tools, advise on good practice, teach, and keep software alive across the lifetime of individual grants.

For research administrators the RSE question is a staffing and budgeting one. Software maintenance is a real and recurring cost that is awkward to fit into time-limited project grants, and institutions differ in whether they provide RSE capacity as a shared service, charge it to individual awards, or leave it to the principal investigator. Titles and job families for RSEs are not standardized, which can complicate recruitment, classification and promotion, and which is one reason the community advocates for recognized career tracks.

Reproducibility depends on software

Computational results can only be checked if the code, its dependencies and the computing environment are available. Software engineering supplies the practices that make this realistic, and they are the practices RSEs promote:

  • Version control for every script and every release, so that a published result can be tied to an exact state of the code.
  • Dependency and environment management through lock files, package managers and containers; see sharing a reproducible conda environment.
  • Workflow management so that multi-step analyses are re-run in a defined order; see Snakemake vs Nextflow.
  • Automated testing and continuous integration, which catch regressions when code, libraries or platforms change.
  • Documentation and code review, which spread understanding beyond a single author; see code review for research software.
  • Archiving of released versions in a repository that issues a persistent identifier.

It helps to be precise about what is being asked. CASRAI’s comparison of reproducibility and replicability distinguishes re-running the same analysis on the same data from confirming a finding with new data, and the dictionary entry on computational reproducibility covers the former, which is the part software practice most directly controls. Journals and funders increasingly ask for a code availability statement, and planning for software is formalized in the software management plan; see the guide to writing a software management plan.

Software citation

If software is a research output that others rely on, it should be citable. The Software Citation Principles, published in 2016 in PeerJ Computer Science by Arfon Smith, Daniel Katz and Kyle Niemeyer on behalf of the FORCE11 Software Citation Working Group, set out the argument that software should be treated as a legitimate, citable product of research. The principles cover importance, credit and attribution, unique identification, persistence, accessibility and specificity; see the dictionary entry on the FORCE11 Software Citation Principles.

In practice, citing software well involves a few concrete steps. Archive each release in a repository that mints a persistent identifier such as a DOI. Include a machine-readable citation file; the Citation File Format (CITATION.cff) is the widely used convention, and GitHub uses it to offer a “cite this repository” option. Cite the specific version used, and cite the software alongside, not instead of, any related paper. For a step-by-step treatment, see how to cite software, code and R packages.

The FAIR Principles for Research Software (FAIR4RS), version 1.0, were released in 2022 by a working group convened jointly by the Research Data Alliance, the Research Software Alliance and FORCE11. They adapt the findable, accessible, interoperable and reusable principles to the particular properties of software, such as executability, composite structure and continuous versioning. See the dictionary entry on FAIR4RS and the guide on how FAIR4RS diverges from FAIR for data.

Publication venues exist for software itself. The Journal of Open Source Software (JOSS) reviews and publishes short papers describing research software; see the entry on JOSS. Other relevant venues include the Journal of Open Research Software and SoftwareX, as well as the main software engineering journals and conferences such as IEEE Transactions on Software Engineering, ACM Transactions on Software Engineering and Methodology, and the International Conference on Software Engineering (ICSE).

Methods, tools and standards

  • Languages and ecosystems — research software is written in Python, R, C and C++, Fortran, Julia and increasingly Rust, among others, with the choice driven by community and performance needs.
  • Collaboration platforms — Git-based hosting for version control, issue tracking and review.
  • Build, packaging and containers — package managers, environment files and container images that capture a runtime environment.
  • Testing and analysis tools — test frameworks, linters, static analyzers and continuous-integration services.
  • High-performance computing — parallel programming, performance profiling and portability across cluster and accelerator architectures.
  • Empirical methods — controlled experiments, case studies, surveys and repository mining.
  • Standards — international standards such as ISO/IEC/IEEE 12207 for software life cycle processes; and for regulated software, domain standards such as IEC 62304 for medical device software.

Who funds software engineering research

In the United States, the National Science Foundation is the principal funder. The Directorate for Computer and Information Science and Engineering funds foundational work, including through the Software and Hardware Foundations (SHF) program; see the entry on NSF CISE. Separately, NSF’s Office of Advanced Cyberinfrastructure administers the Cyberinfrastructure for Sustained Scientific Innovation (CSSI) program, which supports the development of software and data infrastructure for science and succeeded earlier NSF cyberinfrastructure software and data programs. CSSI has funded three classes of award: Elements, Framework Implementations and Transition to Sustainability. Because program solicitations change (when last checked, the CSSI program page listed no open due dates and was awaiting a new solicitation), check the current NSF solicitation before planning a submission.

Other US agencies, including the Department of Energy, the Department of Defense and NASA, fund software engineering research and research-software development where it supports their missions, and NIH institutes fund software development within biomedical informatics and data-science programs. Private foundations have also funded research-software sustainability initiatives. Outside the US, national research councils and the European Union’s framework programmes fund the field, and several countries have dedicated research-software programs.

For applicants, the practical lesson is that funders increasingly expect a plan for the software a project will produce. A software management plan, a named maintenance strategy and a licensing decision belong in the proposal; see the guide on writing a software management plan.

Education and careers

Entry routes include bachelor’s degrees in software engineering or computer science, with many programs ABET-accredited in the US, followed by master’s or PhD study for research and advanced roles. Software engineering doctorates are typically coursework followed by a dissertation based on original empirical or technical work. Professional licensure of software engineers is uncommon compared with civil or mechanical engineering, and employer-recognized experience and portfolios matter more in most roles.

Research-oriented careers include faculty and national-laboratory positions, industrial research labs, and the RSE track in university research-computing groups. Many RSEs hold a PhD in a scientific discipline plus substantial programming experience, while others come from computer science or software engineering degrees; the field does not have a single standard credential.

How software engineering connects to research administration

  • Funding and budgeting — software development and maintenance costs should appear in proposals, and sustaining software beyond an award period needs a plan.
  • Data and software management plans — funders increasingly ask how code will be managed, shared and preserved.
  • Credit and authorship — contributors who write research software deserve recognition, and the roles they play can be described with the CRediT taxonomy, which includes a Software role.
  • Intellectual property and licensing — software raises questions about institutional ownership, open-source licensing and technology transfer that differ from those for data or publications.
  • Research security and export control — some software and technical data are subject to export controls; see the entry on software export controls in the CASRAI dictionary.
  • Integrity and reproducibility policy — code availability requirements, rigor policies and reproducibility audits depend on software practice; see the entry on the reproducibility crisis.
  • Workforce — recruitment, classification and retention of RSEs raise human-resources questions that research offices are often asked to settle.

Frequently asked questions

What is software engineering in simple terms?

Software engineering is the disciplined practice of building and maintaining software so that it works reliably, can be changed safely and can be understood by people other than its author. It covers requirements, design, coding, testing, release and long-term maintenance.

What is the difference between software engineering and computer science?

Computer science studies computation and its theory, including algorithms, languages and systems. Software engineering focuses on the practical methods for producing and maintaining dependable software at scale. The fields overlap and share many courses.

What is a research software engineer?

A research software engineer (RSE) is a professional who combines software engineering skill with knowledge of the research process, building and maintaining the software that research depends on. The term originated in the UK in 2012, and RSE communities now exist in many countries, including the US-RSE association in the United States.

Why does software engineering matter for reproducibility?

Computational results can only be re-checked if the code, dependencies and environment are preserved. Version control, environment management, automated testing and archiving of releases are software engineering practices that make that possible.

How do I cite research software?

Archive the exact version you used in a repository that issues a persistent identifier, cite that identifier and version, and use the project’s citation file if one exists. The FORCE11 Software Citation Principles describe the underlying expectations. See how to cite software, code and R packages.

What are the FAIR principles for research software?

FAIR4RS version 1.0, released in 2022, adapts the findable, accessible, interoperable and reusable principles to software. It was developed by a working group convened by the Research Data Alliance, the Research Software Alliance and FORCE11.

Who funds software engineering research?

In the US, mainly the National Science Foundation, through CISE programs such as Software and Hardware Foundations and through the Office of Advanced Cyberinfrastructure’s CSSI program for scientific software infrastructure, with additional support from mission agencies and foundations.

Do software engineers need a license?

Generally no. Licensure is uncommon in software engineering, unlike civil or mechanical engineering, although accredited degrees and domain-specific certifications can matter in regulated sectors.

Related CASRAI resources

This guide belongs to CASRAI’s series of discipline explainers. See also the guides to computer science, data science, systems engineering, biomedical engineering and biostatistics, and the overview of the branches of science.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about What Is Software Engineering? Research Software, Reproducibility, and Citation

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

Research-admin question? Get an answer that links its sources.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.