Software Open Access

augsig: A Python Package for Physiological Signal Augmentation

Davood Fattahi ,  Xiao Hu

Published Oct. 6, 2026 · Version 0.2.0
When using this resource, please cite:

Fattahi, D., & Hu, X. (2026). augsig: A Python Package for Physiological Signal Augmentation (version 0.2.0). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/5qr1-ka61

Please include the standard citation for PhysioNet:

Pollard, T., Moody, B. E., Lehman, L., Gow, B., Fernandes, C., Xie, C., Johnson, A., Mark, R. G., & Heldt, T. (2026). PhysioNet as a global platform for biomedical research. Nature Health. https://doi.org/10.1038/s44360-026-00096-z. Available from: https://rdcu.be/faatM

Abstract

The augsig package is a modular and extensible Python library for physiological signal augmentation. It provides a set of physiologically consistent augmentation methods for biomedical signals, designed to enhance data diversity, improve generalization of machine learning models, and support controlled signal-level experiments. Its NumPy-based implementation focuses on reproducibility, signal morphology preservation, and extensibility. Augmentation methods include time warping, amplitude modulation and drift, noise injection, and artifact simulation. Time warping and amplitude transformations are realized through Bézier and PCHIP curves, providing smooth, parametric deformations with mathematically guaranteed monotonicity. Each augmentation method exposes multiple configurable parameters, offering fine-grained control over the nature and degree of signal distortion. A multi-recipe pipeline architecture enables arbitrary combinations of transforms to be composed in a single call, supporting diverse augmentation strategies from simple single-transform perturbations to complex compound distortions.


Background

Data augmentation refers to the systematic generation of new data samples by applying controlled perturbations or transformations to existing ones. Rather than collecting additional real-world observations, which is often expensive, time-consuming, or ethically constrained, augmentation constructs plausible variants of available samples to serve a range of purposes. In training, augmented samples expose the model to a broader input distribution, reducing overfitting and improving generalization [1]. At inference, test-time augmentation (TTA) generates multiple variants of each test input and aggregates their predictions, yielding more robust and calibrated outputs [2]. Beyond model training and evaluation, augmentation also supports controlled stress-testing, sensitivity analysis, and simulation of acquisition conditions that are difficult to reproduce experimentally [3].

Physiological signal analysis presents a particularly compelling case for augmentation. Signals such as electrocardiography (ECG) and photoplethysmography (PPG) are time series whose acquisition requires clinical instrumentation and expert annotation, making large labeled datasets difficult and costly to assemble. Furthermore, many clinically relevant conditions, including arrhythmias, hemodynamic instability, and apneic episodes, are inherently rare, producing severe class imbalance that hampers the training of data-hungry deep learning models [4]. At the same time, physiological signals are subject to a wide range of naturally occurring and acquisition-induced variations: heart rate fluctuations alter cycle lengths, respiratory modulation changes signal amplitude, sensor motion introduces baseline drift and burst artifacts, and powerline interference or electromagnetic coupling adds periodic noise [5, 6]. Augmentation methods that faithfully simulate these sources of variation can help models learn robust representations that transfer across subjects, sessions, and devices.

However, augmenting physiological signals imposes constraints that do not apply to images or text. Transformations must preserve signal length, since models are typically trained on fixed-duration windows. They must respect morphological structure, such as the PQRST complex in ECG or the systolic–dicrotic waveform in PPG, to avoid corrupting diagnostically meaningful patterns. Frequency content must remain physiologically plausible, as spectral distortions can introduce artifacts that do not occur in real recordings. These requirements rule out many generic augmentation approaches and motivate the development of domain-aware methods.

Despite a growing literature on physiological signal augmentation [4, 7, 8, 9, 10], existing tools leave key requirements unmet. NeuroKit2 [11] is a comprehensive Python toolbox for ECG, PPG, EDA, and EMG processing with extensive signal preprocessing, feature extraction, and visualization; it provides no dedicated augmentation pipeline, no stochastic variant generation, no SNR control, and no nonlinear warping. tsaug [12] is a general-purpose time-series augmentation library offering noise injection, time warping, and baseline drift; it treats all time series uniformly, providing no spectral color control, no SNR-based scaling, and no domain-specific artifact simulation. audiomentations [13] offers rich audio augmentation including pitch shifting and room simulation; its transforms target audio-domain phenomena and its frequency parameters assume audio sampling conventions, making it unsuitable for physiological workflows. Um et al. [14] established the empirical foundation for wearable sensor augmentation with jitter, scaling, time warping, and magnitude warping; the accompanying code targets inertial sensor data and is not distributed as an installable package. BioAug [15] provides ten augmentation methods for biosignals in NumPy; it offers no SNR targeting, no spectral color control, no burst or powerline artifact simulation, and no monotonicity-guaranteed warping. ecgmentations [16] provides over thirty ECG-specific transforms with a composable pipeline interface; it is restricted to ECG and provides no SNR-based noise control. braindecode [17] includes a mature augmentation module with over twenty transforms and systematic empirical evaluation for EEG and MEG, exemplifying the dedicated tooling available for brain signals; no comparable package has been available for physiological signals.

The augsig package was designed to address these gaps. It provides spectrally shaped noise injection with multiple configurable sample distributions, physiological artifact simulation, nonlinear time warping with monotonicity guarantees, smooth amplitude transforms, a composable multi-recipe pipeline, and sampling-rate-agnostic parameterization. Each transform exposes multiple parameters, enabling fine-grained control suited to diverse signal types, acquisition conditions, and modeling requirements.


Software Description

The augsig package is designed for augmenting one-dimensional physiological signals, built around three guiding principles: modularity, reproducibility, and framework independence. The package is implemented in pure NumPy and SciPy, imposes no dependency on any machine learning framework, and can be integrated into arbitrary signal-processing or model training pipelines regardless of the underlying modeling stack.

The repository is organized as follows:

augsig/
├── augsig/
│   ├── __init__.py       # Re-exports Augment and augment
│   ├── augmenter.py      # Main API: Augment class + augment() function
│   ├── noisifier.py      # Colored noise, burst artifacts, SNR control
│   ├── warper.py         # Time warping, amplitude drift/modulation (Bézier/PCHIP)
│   └── utils.py          # Min-max normalization, Butterworth filtering
├── tests/
│   ├── test.py           # Visual demo script (generates PNG plots)
│   └── test_augsig.py    # pytest suite (137 tests)
├── data/
│   └── sample_ppg.npy    # Sample PPG signal used by the demo and the examples
├── requirements.txt
├── pyproject.toml
├── LICENSE
└── AUTHORS.txt

The package is organized into four focused modules. augmenter.py is the public entry point, defining the augment() function and the Augment class, which orchestrate the full augmentation pipeline. noisifier.py handles noise and artifact generation, providing SNR-controlled additive noise and transient burst artifacts. warper.py implements smooth time-axis and amplitude-axis deformations via Bézier curves and Piecewise Cubic Hermite Interpolating Polynomial (PCHIP) splines, as well as linear baseline drift. utils.py provides shared utilities including min-max normalization and a Butterworth filter wrapper that accepts Nyquist-normalized cutoff frequencies, making the package sampling-rate agnostic.

The sample signal in data/sample_ppg.npy is a real 30-second PPG recording, stored as a 1D array of 1200 samples. It is taken from the BIDMC PPG and Respiration Dataset on PhysioNet [18] (record bidmc01, PLETH channel), resampled from 125 Hz to 40 Hz, and is used by the demo script and the README examples.


Technical Implementation

Augmentation methods for physiological signals can be conceptually divided into two broad categories. The first category introduces non-informative or undesired components into the signal (noise, artifacts, or baseline drift) to expose models to signal degradations that occur in real-world measurements. The second category applies controlled and physiologically plausible variations to the informative structure of the signal, such as subtle changes in timing, amplitude, or morphology, to improve generalizability by simulating natural inter-subject or intra-subject variability.

All methods operate on the raw sample array without requiring knowledge of the sampling rate, as all frequency parameters are expressed in Nyquist-normalized units. Transforms must preserve signal length, respect morphological structure, and maintain physiologically plausible frequency content; these constraints rule out many generic augmentation approaches and motivate the domain-aware design of augsig. Stochastic transforms can be optionally seeded through np.random.default_rng, ensuring exact reproducibility when a fixed seed is provided.

Additive noise supports five spectral colors (white, pink, brown, blue, violet) defined by their power spectral density, providing coarse spectral shaping across the full frequency range. The spectrum can be further refined through an optional Butterworth bandpass filter with configurable cutoff frequencies, constraining noise energy to any desired frequency band. Four sample distributions are available for generating noise samples: Gaussian, uniform, Laplace, and empirical resampling from the input signal or an external pool. Noise amplitude is scaled to a user-specified signal-to-noise ratio (SNR) in dB, allowing precise control over the degree of degradation imposed on the signal.

Time warping, amplitude drift, and amplitude modulation are built on a shared mathematical foundation of Bézier curves and Piecewise Cubic Hermite Interpolating Polynomial (PCHIP) splines. For time warping, the mapping function must be strictly monotone to preserve chronological order; both Bézier and PCHIP formulations guarantee monotonicity by bounding the displacement of interior control points so that no two adjacent knots can cross. Signal length is preserved by construction: the warping curve is anchored at both signal boundaries, mapping the time axis onto itself so the output inherently has the same length as the input. Bézier warping produces globally smooth distortions; PCHIP warping enables sharper, segment-wise deformations. Both formulations share two tuning parameters: the number of control points, which determines the spatial density of deformation, and the perturbation variance, which controls the displacement magnitude. For time warping, the variance is internally capped as a function of the number of control points to preserve the monotonicity guarantee. Amplitude drift and modulation reuse the same curve construction but without the monotonicity constraint, as additive baseline deviations and multiplicative envelopes do not require strictly increasing behavior. Linear baseline drift is also available as a simpler additive trend with configurable slope and intercept.

Burst artifacts simulate short, high-intensity transient disturbances from sensor motion or contact loss, generated by masking spectrally shaped noise (drawn from a Laplace distribution by default to capture heavy-tailed artifact morphology) into a configurable number of randomly placed windows with configurable width, amplitude, and DC offset. The spectral content of the burst noise is additionally shaped by a configurable frequency cutoff, allowing simulation of high-frequency transients (electrode pop) or lower-frequency contact-loss artifacts. Low-frequency noise simulates respiratory or movement-induced baseline oscillations via low-pass filtered white noise. Sinusoidal interference simulates powerline contamination (50/60 Hz or any user-specified frequency) with random phase. Temporal flip and amplitude inversion are deterministic operations that reverse signal orientation along the time or amplitude axis, encouraging models to learn representations invariant to acquisition-level polarity differences.

Transforms are combined through a recipe-based pipeline: the configuration is a dictionary in which each top-level key defines an independent augmentation recipe, specifying which transforms to apply and their parameters. Setting num_copies in a recipe generates multiple stochastic variants from the same configuration. This design supports arbitrary combinations of transforms in a single call, enabling both training-time augmentation and test-time augmentation ensembles without modifying source code. Within a single recipe, transforms are applied sequentially to the same signal, enabling compound augmentations such as a simultaneously warped, drifted, and noisy variant in a single call.

augsig exposes two equivalent interfaces. The function-style API (augment()) is suited for one-shot use, accepting a signal and configuration dictionary and returning an (N, K) array where the first column is the original and each subsequent column is an augmented variant. The class-style API (Augment) binds the configuration and seed at construction time, producing a reusable callable convenient for training loops and dataset classes where the same augmentation policy is applied repeatedly across samples.

For a complete parameter reference and usage examples, readers are referred to the package README. For a detailed methodological account of the augmentation methods, readers are referred to the accompanying publication.


Installation and Requirements

Requirements

  • Python >= 3.8
  • numpy >= 1.21
  • scipy >= 1.7
  • matplotlib >= 3.4 (optional; required only for the visual demo script tests/test.py)

No machine learning framework is required; the package is compatible with any Python environment regardless of whether PyTorch, TensorFlow, or similar libraries are present.

Installation

The recommended installation method is via PyPI:

pip install augsig

Alternatively, the package can be installed from source:

git clone https://github.com/davood-fattahi/augsig.git
cd augsig
pip install .

For environments without package management (e.g., restricted HPC nodes), the augsig/ directory can be copied directly into the project directory and imported without installation, provided numpy and scipy are available:

from augsig import augment, Augment

Usage Notes

augsig exposes two equivalent interfaces. The function-style API (augment(signal, config, seed=None, normalize_output=True)) is suitable for one-shot calls. It accepts a 1D NumPy array of shape (N,) and a nested configuration dictionary, and returns an (N, K) array where the first column is the unmodified original and each subsequent column is one augmented variant. Each top-level key in the configuration dictionary defines one augmentation recipe; a recipe can generate multiple stochastic variants via the num_copies key. Output is min-max normalized to [0, 1] by default; pass normalize_output=False to preserve the original amplitude scale.

The class-style API (Augment(config, seed=None, normalize_output=True)) binds the configuration and seed at construction time, producing a reusable callable. This is convenient for use inside training loops or dataset classes, where the same augmentation policy is applied repeatedly across samples. Omitting the seed produces a different stochastic realization on each call; fixing it yields fully reproducible output.

Inputs carrying singleton axes, such as (N, 1) or (1, N), are squeezed automatically; any other shape raises ValueError. A full configuration key reference covering all supported augmentation parameters is provided in the package README. Readers seeking a detailed methodological account of the augmentation methods are referred to the accompanying publication.


Release Notes

v0.2.0 (current release): first version published on PhysioNet. Available on PyPI (pip install augsig) and on GitHub at https://github.com/davood-fattahi/augsig.

Changes since v0.1.1: adds normalize_output to augment() and Augment; accepts (N, 1) and (1, N) inputs; validates k in the warping functions; fixes resample_pool with an external array; n_bursts=0 now adds no burst noise; clear package-level error for signals too short to filter; pytest suite extended to 137 tests.

v0.1.1 (PyPI only, superseded): initial public release on PyPI and GitHub. Not published on PhysioNet.


Ethics

This project is a software package and does not involve the collection, processing, or distribution of human subjects data. No institutional review board (IRB) approval was required. The sample signal included in the repository (data/sample_ppg.npy) is a 30-second PPG segment drawn from the publicly available BIDMC PPG and Respiration Dataset (PhysioNet, record bidmc01, PLETH channel, resampled from 125 Hz to 40 Hz), which is released under the Open Data Commons Attribution License (ODC-BY) and contains no personally identifiable information.

The primary benefit of augsig is that it enables researchers to expand limited physiological signal datasets without requiring additional data collection from human participants, thereby reducing both cost and ethical burden. It may also help address class imbalance in rare clinical conditions, potentially improving the fairness and robustness of machine learning models.

The principal risk is misuse: augmented signals are synthetic variants and should not substitute for real clinical recordings in safety-critical validation or regulatory submissions. Users are responsible for ensuring that augmentation parameters produce physiologically plausible variants appropriate to their application, and that models trained on augmented data are validated against real-world recordings before clinical deployment.


Acknowledgements

This work was conducted at the Nell Hodgson Woodruff School of Nursing and the Center for Data Science, Emory University. This work was partially supported by the National Institutes of Health through grant number R01HL166233.


Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. [1] C. Shorten and T. M. Khoshgoftaar, "A survey on Image Data Augmentation for Deep Learning," Journal of Big Data, vol. 6, no. 1, p. 60, 2019. doi:10.1186/s40537-019-0197-0
  2. [2] M. Kimura, "Understanding Test-Time Augmentation," in Neural Information Processing, T. Mantoro et al., Eds. Cham: Springer International Publishing, 2021, pp. 558–569.
  3. [3] N. Nemati, "Comparative Analysis of Data Augmentation for Clinical ECG Classification with STAR," medRxiv, 2025.
  4. [4] P. Cao et al., "A novel data augmentation method to enhance deep neural networks for detection of atrial fibrillation," Biomedical Signal Processing and Control, vol. 56, p. 101675, 2020. doi:10.1016/j.bspc.2019.101675
  5. [5] H. Lee, H. Chung, and J. Lee, "Motion Artifact Cancellation in Wearable Photoplethysmography Using Gyroscope," IEEE Sensors Journal, vol. 19, no. 3, pp. 1166–1175, 2019. doi:10.1109/jsen.2018.2879970
  6. [6] T. Pereira et al., "Deep learning approaches for plethysmography signal quality assessment in the presence of atrial fibrillation," Physiological Measurement, vol. 40, no. 12, p. 125002, 2019. doi:10.1088/1361-6579/ab5b84
  7. [7] M. Guhdar, R. J. Mstafa, and A. O. Mohammed, "A novel data augmentation strategy for robust deep learning classification of biomedical time-series data: Application to ECG and EEG analysis," arXiv:2507.12645, 2025.
  8. [8] J. An, R. E. Gregg, and S. Borhani, "Effective Data Augmentation, Filters, and Automation Techniques for Automatic 12-Lead ECG Classification Using Deep Residual Neural Networks," in Proc. 44th Annual Int. Conf. IEEE Engineering in Medicine & Biology Society (EMBC), 2022, pp. 1283–1287. doi:10.1109/EMBC48229.2022.9871654
  9. [9] P. Guo, H. Yang, and A. Sano, "Empirical Study of Mix-based Data Augmentation Methods in Physiological Time Series Data," in Proc. IEEE 11th Int. Conf. Healthcare Informatics (ICHI), 2023, pp. 206–213.
  10. [10] M. F. Safdar, P. Pałka, A. A. Faresi, and R. M. Nowak, "Optimizing Electrocardiogram Signal Augmentation for Realistic Synthetic Data in Deep Learning Model," in Proc. Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA), 2024, pp. 54–59. doi:10.23919/SPA61993.2024.10715629
  11. [11] D. Makowski et al., "NeuroKit2: A Python toolbox for neurophysiological signal processing," Behavior Research Methods, vol. 53, no. 4, pp. 1689–1696, 2021. doi:10.3758/s13428-020-01516-y
  12. [12] Arundo Analytics, "tsaug: A Python package for time series augmentation," 2020. Available: https://github.com/arundo/tsaug
  13. [13] I. Jordal, "audiomentations: A Python library for audio data augmentation," 2019. Available: https://github.com/iver56/audiomentations
  14. [14] T. T. Um et al., "Data augmentation of wearable sensor data for Parkinson's disease monitoring using convolutional neural networks," in Proc. 19th ACM Int. Conf. Multimodal Interaction (ICMI), Glasgow, UK, 2017, pp. 216–220. doi:10.1145/3136755.3136817
  15. [15] P. Chen, "BioAug: A toolbox for biosignal augmentation," 2023. Available: https://github.com/peijii/BioAug
  16. [16] R. Epifanov, "ecgmentations: ECG augmentation library inspired by Albumentations," 2023. Available: https://github.com/rostepifanov/ecgmentations
  17. [17] R. T. Schirrmeister et al., "Braindecode: Open-source software for decoding raw electrophysiological brain signals," Zenodo, 2023. doi:10.5281/zenodo.17699192. Available: https://github.com/braindecode/braindecode
  18. [18] M. A. F. Pimentel, A. E. W. Johnson, P. H. Charlton, D. Birrenkott, P. J. Watkinson, L. Tarassenko, and D. A. Clifton, "Toward a Robust Estimation of Respiratory Rate From Pulse Oximeters," IEEE Transactions on Biomedical Engineering, vol. 64, no. 8, pp. 1914-1923, 2017. doi:10.1109/TBME.2016.2613124. Dataset: BIDMC PPG and Respiration Dataset, PhysioNet. doi:10.13026/C2208R

Files

Total uncompressed size: 80.5 KB.

Access the files
Folder Navigation: <base>
Name Size Modified
augsig
data
tests
AUTHORS.txt (download) 636 B 2026-09-22
LICENSE (download) 1.5 KB 2026-09-22
LICENSE.txt (download) 1.5 KB 2026-09-29
README.md (download) 15.7 KB 2026-09-22
SHA256SUMS.txt (download) 1.1 KB 2026-09-29
pyproject.toml (download) 1.4 KB 2026-09-22
requirements.txt (download) 99 B 2026-09-22