summary() describes a tbl_now the way a nowcaster needs it described:
how many cases arrive on each of the object's time axes, how long they take
to get there, how sparse the series is, what fraction of the data is
censored or still pending, and how far the object reaches.
Every block of the summary is also available on its own – see
nowcast_summary_components – and summary() is exactly the
dplyr::bind_rows() of those pieces.
Usage
# S3 method for class 'tbl_now'
summary(object, ..., by_strata = NULL, strata = NULL, growth_k = 7)Arguments
- object
A
tbl_nowobject.- ...
Unused, for compatibility with the
base::summary()generic.- by_strata
Logical. Add one set of rows per stratum on top of the pooled (
"all") rows. Defaults toTRUEwhen the object has strata.- strata
Character vector of columns to stratify by. Defaults to
get_strata(object).- growth_k
Number of delays for the cumulative growth rows.
Details
The date grids. "Cases per event date" is a statement about a calendar,
not about the rows present in the data, so each axis is completed to a full
grid running from the earliest observed date on that axis to get_now(),
stepping by that axis's units. Dates with no rows count as zeros. This is
what makes prop_zero and the zero-run lengths meaningful, and it is why a
line list – which cannot represent a zero – is summarised correctly here.
Not-yet-observed cells are dropped. An NA count means the cell has not
been observed yet, unlike a 0, which was observed and was zero. Such rows
carry no cases, so they are excluded rather than allowed to turn every total
they touch into NA. How many were dropped is reported as the
"unobserved_cells" coverage row.
The grid is global: when by_strata = TRUE every stratum is summarised
on the same grid, so a stratum whose cases start late genuinely shows the
leading zeros. Otherwise the strata would not be comparable.
Count-cumulative data gets no delay rows. A cumulative total is not
additive across delays, so a case-weighted delay distribution would be
meaningless. The "growth" rows take their place, describing how each event
date's total grows from one delay to the next. Call
to_count(x, to = "count-incidence") first if you want the delay
distribution – and note that de-accumulating can produce negative
increments.
Note
Quantiles are inverse-ECDF (type 1), not stats::quantile()'s default.
q50 is the smallest value whose cumulative weight reaches 0.5, which for
an even number of observations is the upper of the two middle values rather
than their average. This is deliberate: it is the same estimator
autoplot.tbl_now() and diagnose_drift() use for the delay quantiles
they draw, so the numbers in this table match the numbers in the plots. It
also always returns a value that was actually observed, which a half-case
delay is not.
The columns
Every function in this family returns the same schema, so results can be
stacked with dplyr::bind_rows() and filtered with dplyr::filter().
componentWhich block the row belongs to:
"cases","delay","zero_run","composition","growth"or"coverage".quantityWhat the row describes, including the category for the compositional rows (
"revision_type = confirmed").stratumWhich subset of the data the row describes:
"all"for the pooled rows, or the stratum label otherwise.nHow many things of the block's own kind the row counts. It is never a case count, and what it counts changes with the block: dates on the grid for
"cases", runs of zeros for"zero_run", and (event date, report date) cells for"delay"and"composition".totalHow many cases are behind the row: records for a line list, the sum of the case-count column otherwise. So
nandtotalanswer different questions and are equal only when every cell holds exactly one case. In the"composition"block, for instance,nis how many cells carry that category andtotalis how many cases do – andpropis computed fromtotal, the cases.mean,sdMean and standard deviation. For the case-weighted rows these are the weighted versions, equal to what you would get by expanding the counts to one row per case.
min,q25,q50,q75,q90,maxQuantiles. See the note below on which estimator is used.
prop_zeroProportion of dates on the grid that are exactly zero.
propProportion of cases in this category (compositional rows).
valueA single scalar that is not a distribution: a gap or an occupancy. The
"growth"rows are distributions over event dates, so they populatemean/sd/the quantiles instead and leavevalueempty.date_min,date_maxDate range. Present only when the result contains
"coverage"rows.unobserved_cellsA
"coverage"row counting theNA-count rows excluded as not yet observed.
See also
nowcast_summary_components for the individual blocks.
Examples
data(denguedat)
# The last five years. The full twenty-year series gives the same shape
# of answer, it just takes longer to compute.
recent <- denguedat[denguedat$onset_week >= as.Date("2006-01-01"), ]
ndata <- tbl_now(recent,
event_date = "onset_week",
report_date = "report_week",
strata = "gender",
verbose = FALSE
)
# The whole summary: one row per quantity, per stratum.
overview <- summary(ndata)
overview
#> ── Summary of a <tbl_now> ──────────────────────────────────────────────────────
#> 46 rows in 5 components; strata: "Female" and "Male".
#>
#> cases
#> n = dates on the grid; total = cases
#> quantity stratum n total mean sd min q25 q50 q75 q90 max
#> <chr> <chr> <int> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 per_event… all 260 14135 54.4 73.2 0 11 25 65 139 358
#> 2 per_event… Female 260 6998 26.9 36.4 0 6 12 31 71 189
#> 3 per_event… Male 260 7137 27.4 37.2 0 5 13 32 71 176
#> 4 per_repor… all 259 14135 54.6 75.0 0 10 25 65 142 420
#> 5 per_repor… Female 259 6998 27.0 37.3 0 5 12 33 73 217
#> 6 per_repor… Male 259 7137 27.6 38.0 0 6 14 33 70 203
#> # ℹ 1 more variable: prop_zero <dbl>
#>
#> zero_run
#> n = runs of consecutive zero dates; total = zero dates in those runs
#> quantity stratum n total mean sd min q25 q50 q75 q90 max
#> <chr> <chr> <int> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 event_da… all 1 3 3 NA 3 3 3 3 3 3
#> 2 event_da… Female 5 8 1.6 0.894 1 1 1 2 3 3
#> 3 event_da… Male 5 7 1.4 0.894 1 1 1 1 3 3
#> 4 report_d… all 2 2 1 0 1 1 1 1 1 1
#> 5 report_d… Female 8 9 1.12 0.354 1 1 1 1 2 2
#> 6 report_d… Male 7 8 1.14 0.378 1 1 1 1 2 2
#>
#> composition
#> n = (event, report) cells in the category; total = cases in the category
#> quantity n total prop
#> <chr> <int> <dbl> <dbl>
#> 1 strata = Female 842 6998 0.495
#> 2 strata = Male 831 7137 0.505
#>
#> coverage
#> n = cells, or distinct dates on a date row; total = cases
#> quantity stratum n total date_min date_max
#> <chr> <chr> <int> <dbl> <date> <date>
#> 1 total_cases all 1673 14135 NA NA
#> 2 event_date all 257 14135 2006-01-02 2010-11-29
#> 3 report_date all 257 14135 2006-01-09 2010-12-20
#> 4 total_cases Female 842 6998 NA NA
#> 5 event_date Female 252 6998 2006-01-02 2010-11-29
#> 6 report_date Female 250 6998 2006-01-09 2010-12-20
#> 7 total_cases Male 831 7137 NA NA
#> 8 event_date Male 253 7137 2006-01-02 2010-11-29
#> 9 report_date Male 251 7137 2006-01-09 2010-12-13
#> 10 now all NA NA 2010-12-20 2010-12-20
#> ℹ 19 more rows.
#>
#> delay
#> n = (event, report) cells; total = cases
#> quantity stratum n total mean sd min q25 q50 q75 q90 max
#> <chr> <chr> <int> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 event_to_… all 1673 14135 1.81 1.06 0 1 2 2 3 26
#> 2 event_to_… Female 842 6998 1.82 1.07 0 1 2 2 3 15
#> 3 event_to_… Male 831 7137 1.80 1.06 0 1 2 2 3 26
#>
#> ℹ Use `dplyr::filter()` or `tibble::as_tibble()` for the full schema.
# It is an ordinary tibble, so pick out the block you want.
overview |> dplyr::filter(component == "delay")
#> ── Summary of a <tbl_now> ──────────────────────────────────────────────────────
#> 3 rows in 1 component; strata: "Female" and "Male".
#>
#> delay
#> n = (event, report) cells; total = cases
#> quantity stratum n total mean sd min q25 q50 q75 q90 max
#> <chr> <chr> <int> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 event_to_… all 1673 14135 1.81 1.06 0 1 2 2 3 26
#> 2 event_to_… Female 842 6998 1.82 1.07 0 1 2 2 3 15
#> 3 event_to_… Male 831 7137 1.80 1.06 0 1 2 2 3 26
#>
#> ℹ Use `dplyr::filter()` or `tibble::as_tibble()` for the full schema.
# `n` and `total` are different questions. In the compositional block `n`
# counts the event-report cells carrying the category and `total` counts
# the cases in them.
overview |>
dplyr::filter(component == "composition") |>
dplyr::select(quantity, stratum, n, total, prop)
#> # A tibble: 2 × 5
#> quantity stratum n total prop
#> <chr> <chr> <int> <dbl> <dbl>
#> 1 strata = Female all 842 6998 0.495
#> 2 strata = Male all 831 7137 0.505
# Pooled rows only, ignoring the strata.
summary(ndata, by_strata = FALSE)
#> ── Summary of a <tbl_now> ──────────────────────────────────────────────────────
#> 16 rows in 4 components.
#>
#> cases
#> n = dates on the grid; total = cases
#> quantity n total mean sd min q25 q50 q75 q90 max prop_zero
#> <chr> <int> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 per_eve… 260 14135 54.4 73.2 0 11 25 65 139 358 0.0115
#> 2 per_rep… 259 14135 54.6 75.0 0 10 25 65 142 420 0.00772
#>
#> zero_run
#> n = runs of consecutive zero dates; total = zero dates in those runs
#> quantity n total mean sd min q25 q50 q75 q90 max
#> <chr> <int> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 event_date 1 3 3 NA 3 3 3 3 3 3
#> 2 report_date 2 2 1 0 1 1 1 1 1 1
#>
#> coverage
#> n = cells, or distinct dates on a date row; total = cases
#> quantity n total value date_min date_max
#> <chr> <int> <dbl> <dbl> <date> <date>
#> 1 total_cases 1673 14135 NA NA NA
#> 2 event_date 257 14135 NA 2006-01-02 2010-11-29
#> 3 report_date 257 14135 NA 2006-01-09 2010-12-20
#> 4 now NA NA NA 2010-12-20 2010-12-20
#> 5 unobserved_cells 0 NA NA NA NA
#> 6 max_delay NA NA 26 NA NA
#> 7 triangle_cells_observed 1034 NA NA NA NA
#> 8 triangle_cells_possible 6669 NA NA NA NA
#> 9 triangle_occupancy NA NA 0.155 NA NA
#> 10 now_gap_event NA NA 3 NA NA
#> ℹ 1 more row.
#>
#> delay
#> n = (event, report) cells; total = cases
#> quantity n total mean sd min q25 q50 q75 q90 max
#> <chr> <int> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 event_to_report 1673 14135 1.81 1.06 0 1 2 2 3 26
#>
#> ℹ Use `dplyr::filter()` or `tibble::as_tibble()` for the full schema.
