Handles three types of timestamp trouble common in time-series:
Drifted timestamps - a reading scheduled for 13:05 was logged at 13:06 instead. For
unit \%in\% c("sec","min","hour","day"), each observed timestamp is snapped to the nearest point on the regular grid, provided it falls withintoleranceof that point. (Not applicable to"month"/"quarter"/"year", whose calendar steps have no fixed duration - those units always use exact matching, as before.)Missing rows - timestamps that are absent entirely. These rows are inserted with
NAfor all value columns.Missing values - rows that exist but have
NAin some columns.
After filling gaps the result is ready to pass into
missing_analysis() and handle_missing().
Usage
fill_time_gaps(
data,
time_col,
n = 1,
unit = c("sec", "min", "hour", "day", "month", "quarter", "year"),
tolerance = NULL,
format = NULL,
verbose = TRUE
)Arguments
- data
A
data.frame,tibble, ortsibblewith a timestamp column, sorted ascending. tsibble and tibble are converted todata.frameautomatically.- time_col
Character. Name of the timestamp column (must be
DateorPOSIXct). Usecombine_datetime()first if date and time are stored in separate columns.- n
Numeric. Number of units per step. e.g.
1,3,15.- unit
Character. One of
"sec","min","hour","day","month","quarter","year"."month"- calendar months (1, 2, 3, ...)"quarter"- calendar quarters of 3 months each (Q1, Q2, Q3, Q4)
Note:
"month"and"quarter"require the timestamp column to beDateand each value to fall on the first day of the month (e.g.2024-01-01,2024-04-01).- tolerance
Numeric seconds, or
NULL. Only used whenunit \%in\% c("sec","min","hour","day"). Maximum distance an observed timestamp may sit from its nearest grid point and still be snapped to it.NULL(default) uses half the step size (n/unit), the widest tolerance that still avoids ambiguity between two adjacent grid points. Set to0to disable snapping and fall back to the legacy exact-match behavior (a timestamp must land exactly on the grid to count; otherwise it is left as-is and its grid slot is reported as a gap). Ignored (with a warning if non-NULL) for"month","quarter", and"year".- format
Optional. Used only when
time_colischaracter(e.g. read via baseread.csv(), which never auto-detects dates). Astrptime()-style format string (or vector of candidates to try in order), e.g."\%m/\%d/\%Y". WhenNULL(default), a set of common formats is tried automatically and the first one that parses every non-missing value without introducing newNAs is used. If none match,fill_time_gaps()aborts with guidance on how to convert the column yourself.- verbose
Logical (default
TRUE).
Value
An invisible list with elements:
dataComplete
data.framewith no missing timestamps. Inserted rows haveNAfor all value columns.n_gapsNumber of timestamp rows inserted.
n_na_beforeTotal NA cells before gap-filling.
n_na_afterTotal NA cells after gap-filling (includes inserted rows).
gap_timestampsVector of timestamps that were inserted.
time_colName of the timestamp column.
freqFrequency string describing the step used.
n_snappedNumber of observed timestamps snapped onto the grid (
0for calendar units or whentolerance = 0).n_dropped_duplicateRows dropped because a closer observation already claimed the same grid slot.
n_dropped_toleranceRows dropped because they were too far from any grid slot.
Examples
# Sensor drift: 13:06 should have been 13:05
df <- data.frame(
datetime = as.POSIXct(c("2024-01-01 13:00:00",
"2024-01-01 13:06:00",
"2024-01-01 13:10:00")),
vpd = c(1.2, 1.4, 1.5)
)
result <- fill_time_gaps(df, time_col = "datetime", n = 5, unit = "min")
#>
#> ============================================================
#> STEP : Check and Fill Missing Timestamps
#> ============================================================
#>
#> Frequency : every 5 minute(s)
#> Tolerance : +/- 150 sec
#> Expected : 3 rows
#> Found : 3 rows
#> Gaps : 0 rows inserted
#> Snapped : 1 timestamp(s) moved onto the grid
#> Dropped : 0 duplicate(s), 0 out-of-tolerance
#>
#> NA before : 0
#> NA after : 0
#>
#> [OK] No gaps found
#>
#> >> Next: ts_preprocess()
#>
result$data
#> datetime vpd
#> 1 2024-01-01 13:00:00 1.2
#> 2 2024-01-01 13:05:00 1.4
#> 3 2024-01-01 13:10:00 1.5
# Monthly data - missing February (calendar unit: exact match only)
df_m <- data.frame(
date = as.Date(c("2024-01-01", "2024-03-01", "2024-04-01")),
sales = c(100, 130, 150)
)
result_m <- fill_time_gaps(df_m, time_col = "date", n = 1, unit = "month")
#>
#> ============================================================
#> STEP : Check and Fill Missing Timestamps
#> ============================================================
#>
#> Frequency : every 1 month(s)
#> Expected : 4 rows
#> Found : 3 rows
#> Gaps : 1 rows inserted
#>
#> NA before : 0
#> NA after : 1
#> └─ inserted rows : 1 x 1 cols = 1 (filled in next step)
#> └─ pre-existing : 0
#>
#> ------------------------------------------------------------
#> Inserted timestamps (1 of 1)
#> ------------------------------------------------------------
#> [1] "2024-02-01"
#>
#> >> Next: ts_preprocess(gap$data, ...)
#>