Time Series Features in the feasts R Package

A reference guide to feature-extraction functions for tidy time series

Author

Prepared with Claude

Published

28 September 2026

1 Overview

feasts (Feature Extraction And Statistics for Time Series) is part of the tidyverts ecosystem and works with tsibble objects. Its feature functions condense a time series into a small number of interpretable numbers describing things like trend strength, seasonality, autocorrelation structure, and irregularity. These are typically applied with fabletools::features():

Code
library(feasts)
library(tsibble)
library(dplyr)

tourism |>
  features(Trips, feature_set(pkgs = "feasts"))

feature_set(pkgs = "feasts") collects every registered feature function in the package, but each one can also be called individually, as shown throughout this report.

The functions fall into a few natural groups:

  1. Decomposition-based features — derived from an STL decomposition (trend/seasonal strength, spikiness, etc.)
  2. Autocorrelation-based features — from the ACF and PACF of the series and its differences
  3. Spectral and entropy features — how “predictable” or noise-like a series is
  4. Distributional / shape features — flat spots, crossing points, intermittency
  5. Window-based features — lumpiness, stability, and structural shifts over sub-windows
  6. Statistical test features — unit roots, stationarity, heteroscedasticity, cointegration

2 STL decomposition–based features

2.1 feat_stl()

Code
feat_stl(x, .period, s.window = 11, ...)

Decomposes the series into trend T_t, seasonal S_t, and remainder R_t using STL, then computes a set of descriptive statistics from those components. The main quantities are:

Feature Definition Interpretation
trend_strength \max\left(0,\ 1 - \dfrac{\mathrm{Var}(R_t)}{\mathrm{Var}(T_t + R_t)}\right) Close to 1 → strong, dominant trend
seasonal_strength \max\left(0,\ 1 - \dfrac{\mathrm{Var}(R_t)}{\mathrm{Var}(S_t + R_t)}\right) Close to 1 → strong, dominant seasonality
seasonal_peak Season (e.g. month/quarter) of the maximum of the seasonal component Timing of the seasonal high point
seasonal_trough Season of the minimum of the seasonal component Timing of the seasonal low point
spikiness Variance of the leave-one-out variances of R_t Large value → remainder dominated by a few big spikes
linearity Coefficient on a linear term from an orthogonal-polynomial regression fit to the trend Positive → strongly increasing/decreasing trend
curvature Coefficient on a quadratic term from the same regression Large magnitude → trend bends noticeably
stl_e_acf1 First-lag autocorrelation of the remainder High value → remainder is not white noise
stl_e_acf10 Sum of squares of the first 10 remainder autocorrelations High value → substantial leftover structure in the remainder

Spikiness in detail. For remainder series e_t of length n with sample variance \mathrm{Var}(e), the leave-one-out variance for observation i is computed without refitting:

d_i = (e_i - \bar e)^2, \qquad \mathrm{varloo}_i = \frac{\mathrm{Var}(e)(n-1) - d_i}{n-2}

Spikiness is then \mathrm{Var}(\mathrm{varloo}_1, \dots, \mathrm{varloo}_n). If no single point dominates the remainder, all leave-one-out variances are similar and spikiness is near zero; a few extreme spikes make the leave-one-out variances disperse widely, inflating this feature.

Code
tourism |> features(Trips, feat_stl)

3 Autocorrelation-based features

3.1 feat_acf()

Code
feat_acf(x, .period = 1, lag_max = NULL, ...)

Returns 6 (or 7, if seasonal) values built from the autocorrelation function of the original series, its first difference, and its second difference:

  • First-lag autocorrelation and the sum of squares of the first ten autocorrelations, computed for the original series, the once-differenced series, and the twice-differenced series
  • For seasonal data (.period > 1), the autocorrelation at the first seasonal lag is also included

High sums of squared autocorrelations indicate a series with strong short-term dependence; comparing the values across the original/differenced/twice-differenced versions helps identify how much differencing is needed to reach approximate white noise.

3.2 feat_pacf()

Code
feat_pacf(x, .period = 1, lag_max = NULL, ...)

The partial-autocorrelation analogue of feat_acf(): the sum of squares of the first 5 partial autocorrelations for the original, once-differenced, and twice-differenced series, plus the seasonal partial autocorrelation for seasonal data.

3.3 ACF(), PACF(), CCF()

Not single-number features but tidy computations of the (partial) autocorrelation or cross-correlation function, returned as a tbl_cf object that can be plotted with autoplot(). Useful for visual inspection alongside the numeric ACF/PACF-based features above.


4 Spectral and entropy features

4.1 feat_spectral()

Code
feat_spectral(x, .period = 1, ...)

Computes the spectral entropy of the series — the Shannon entropy of its (AR-estimated, Burg-method) normalized spectral density f_x(\lambda):

H_s(x_t) = -\int_{-\pi}^{\pi} f_x(\lambda) \log f_x(\lambda) \, d\lambda, \qquad \int_{-\pi}^{\pi} f_x(\lambda)\, d\lambda = 1

Values close to 0 indicate a series that is easy to forecast (its spectral power is concentrated, e.g. a strong single frequency or trend); values closer to 1 indicate a series closer to white noise, where power is spread evenly across frequencies.

4.2 coef_hurst()

Code
coef_hurst(x)

Estimates the Hurst coefficient, reflecting the degree of long-range dependence / fractional differencing in the series. Values above 0.5 suggest persistent, trending behaviour; values below 0.5 suggest anti-persistent, mean-reverting behaviour.


5 Shape, intermittency, and irregularity features

5.1 feat_intermittent()

Code
feat_intermittent(x)

Designed for intermittent-demand series (many zeros interspersed with non-zero values):

Feature Meaning
zero_run_mean Average interval length between non-zero observations
nonzero_squared_cv Squared coefficient of variation of the non-zero observations
zero_start_prop Proportion of the series that starts with zeros
zero_end_prop Proportion of the series that ends with zeros

5.2 longest_flat_spot()

Code
longest_flat_spot(x)

Splits the range of the series into 10 equal-width bins and returns the longest run of consecutive observations sitting in the same bin — a proxy for periods where the series barely changes.

5.3 n_crossing_points()

Code
n_crossing_points(x)

Counts how many times the series crosses its own median. Frequent crossings suggest choppier, more mean-reverting behaviour; few crossings suggest sustained runs above or below the median.


6 Window-based features

6.1 var_tiled_mean() and var_tiled_var()

Code
var_tiled_mean(x, .size = NULL, .period = 1)   # "stability"
var_tiled_var(x, .size = NULL, .period = 1)    # "lumpiness"

Both split the series into non-overlapping (“tiled”) windows and summarise each window with its mean or its variance:

  • Stability — the variance of the tile means. High values indicate the local level shifts a lot over time.
  • Lumpiness — the variance of the tile variances. High values indicate the series’ volatility itself changes over time (heteroscedasticity across chunks of the series).

6.2 shift_level_max(), shift_var_max(), shift_kl_max()

Code
shift_level_max(x, .size = NULL, .period = 1)
shift_var_max(x, .size = NULL, .period = 1)
shift_kl_max(x, .size = NULL, .period = 1)

Slide a window across the series and compare consecutive windows, returning both the size of the largest change and the time index at which it occurs:

  • shift_level_max — largest shift in the mean between adjacent windows (level/step changes)
  • shift_var_max — largest shift in the variance between adjacent windows (volatility changes)
  • shift_kl_max — largest shift in Kullback–Leibler divergence between adjacent windows (distributional changes more general than mean/variance alone)

7 Statistical-test features

These wrap classical hypothesis tests so their statistics and p-values can be used as machine-learning-style features.

7.1 Unit root and stationarity

Code
unitroot_kpss(x, type = c("mu", "tau"), lags = c("short", "long", "nil"))
unitroot_pp(x, type = c("Z-tau", "Z-alpha"), model = c("constant", "trend"))
unitroot_ndiffs(x, alpha = 0.05, ...)
unitroot_nsdiffs(x, alpha = 0.05, ...)
  • unitroot_kpss — KPSS test statistic; large values point to non-stationarity (rejects the null of stationarity)
  • unitroot_pp — Phillips–Perron test statistic for a unit root
  • unitroot_ndiffs — minimum number of ordinary differences needed for stationarity, based on repeatedly applying a unit-root test (KPSS by default)
  • unitroot_nsdiffs — minimum number of seasonal differences needed, based by default on whether STL seasonal strength exceeds 0.64

7.2 Portmanteau (autocorrelation) tests

Code
ljung_box(x, lag = 1, dof = 0)
box_pierce(x, lag = 1, dof = 0)

Test the null hypothesis that the series (often a residual series) is independently distributed, i.e. shows no autocorrelation up to the specified lag. Both return a statistic and a p-value; Ljung–Box is the small-sample-corrected version of Box–Pierce.

7.3 Heteroscedasticity

Code
stat_arch_lm(x, lags = 12, demean = TRUE)

Computes Engle’s ARCH LM statistic — the R^2 from an autoregression of the squared series — as a measure of autoregressive conditional heteroscedasticity (volatility clustering).

7.4 Cointegration

Code
cointegration_johansen(x, ...)
cointegration_phillips_ouliaris(x, ...)
  • cointegration_johansen — Johansen procedure (trace/eigenvalue statistics) for testing cointegration among multiple series, via urca::ca.jo()
  • cointegration_phillips_ouliaris — Phillips–Ouliaris residual-based cointegration test statistic and approximate p-value, via urca::ca.po()

7.5 Box–Cox transformation selection

Code
guerrero(x, lower = -0.9, upper = 2, .period = 2L)

Not a descriptive feature but a transformation-parameter selector: chooses the Box–Cox \lambda that minimises the coefficient of variation across subseries of the data (Guerrero’s method).


8 Putting it together

A typical workflow computes the full built-in feature set, then explores it — e.g. with PCA — to find unusual or clustered series:

Code
library(feasts)
library(tsibble)
library(dplyr)

feat_tbl <- tourism |>
  features(Trips, feature_set(pkgs = "feasts"))

feat_tbl

Individual functions can also be combined manually for a custom, lighter-weight feature set:

Code
tourism |>
  features(Trips, list(
    feat_stl,
    coef_hurst,
    feat_spectral,
    unitroot_kpss
  ))

9 References

  • Hyndman, R.J., Wang, E., & Laptev, N. (2015). Large-scale unusual time series detection. IEEE ICDM Workshops.
  • Kang, Y., Hyndman, R.J., & Smith-Miles, K. (2017). Visualising forecasting algorithm performance using time series instance space. International Journal of Forecasting.
  • Hyndman, R.J. & Athanasopoulos, G. Forecasting: Principles and Practice, “Time series features” chapter — https://otexts.com/fpp3/features.html
  • feasts package documentation — https://feasts.tidyverts.org/