emoji_trend() counts emoji per time period and returns the long table
that plots directly: one row per (period, emoji) over the periods it
returns, including the ones in which a given emoji is absent, so a trend
line does not silently skip its zeros.
Usage
emoji_trend(
data,
text,
time,
by = "month",
top_n = 20,
measure = c("n", "share")
)Arguments
- data
A data frame or tibble containing a text column. Grouped data frames are accepted. The verbs that work a row at a time (adding columns, or keeping and expanding rows) carry the grouping through to their result, as
dplyr::mutate()anddplyr::filter()do. The verbs that pool across rows – the counts, the co-occurrence edge lists, the time series – warn that they ignore the grouping and return one corpus-wide answer.- text
The text column to scan, supplied unquoted. Any atomic column is accepted and read as character, so a
factorworks and a numeric,Dateor logical one simply contains no emoji. A list column – or a data-frame column – is refused rather than coerced, because coercing one deparses it and the emoji found would be in the code rather than in your data. What counts as an emoji is the same in every verb; see the Detection section of tidyEmoji for the one case that surprises people, code points that are emoji only when they carryU+FE0F.- time
Unquoted column of dates or date-times (
Date,POSIXct, or character in"YYYY-MM-DD"form). A date-time is bucketed by the calendar day it displays as in its own timezone, not by its UTC day: an emoji posted at 23:30 New York time belongs to that day, not to the next one."Its own timezone" means the column's
tzoneattribute. APOSIXctcreated without one – which is whatas.POSIXct("2024-01-01 23:30")and most CSV readers give you – has no timezone of its own, so R displays it in the session's, and the buckets follow. The same column then gives hour 23 on one machine and hour 4 on another. Tag the column (as.POSIXct(x, tz = "UTC"), orlubridate::force_tz()) if the result has to be reproducible; aDatecolumn is immune either way.A character column must lead with a four-digit year:
"2024-01-01"or"2024/01/01", with one- or two-digit month and day, and any trailing time ignored. Values that do not parse warn and are dropped – but a column in which nothing reads as a date is an error rather than a column ofNA, since there would be no time axis left. Note that"01/02/2024"is in the second group: convert a column written that way withas.Date()and its ownformatfirst.- by
Period length:
"day","week"(starting Monday),"month"(default),"quarter"or"year".- top_n
Number of emoji to follow, ranked by
measureover the whole corpus.NULLkeeps every emoji. Default20. When a tie straddles the cut the glyph decides which emoji fall inside it, in the C locale, as intop_n_emojis(); a corpus with fewer emoji than this returns every one of them rather than padding, and0returns no rows at all.- measure
Statistic used to rank emoji for
top_nand to order the rows within a period:"n"(default) or"share".
Value
A tibble with columns .period (a Date, the start of the period),
emoji, name, n and share, sorted by .period, then by measure
descending, then by the glyph, so the order is fully determined. That
last key matters: within a period the zeros this verb fills in all tie on
both of the others.
Details
Which periods appear. The grid is complete over the observed periods,
and "observed" means a period holding at least one emoji. A period whose
rows carry no emoji at all does not appear, and neither does a gap in the
calendar: emoji_trend() never invents a period. So the zeros it fills in
are the ones within the periods it returns, not a continuous time axis.
Pass the result through tidyr::complete() against a calendar sequence if
you need the empty periods too.
The three time verbs answer this differently, on purpose, and it is worth
knowing which you are getting before joining two of them on .period:
emoji_trend()– periods containing at least one emoji.emoji_turnover()– every period containing at least one dated row, including emoji-free ones, which reportn_types = 0.emoji_seasonality()– every level of the cycle unconditionally, whether or not the data reaches it.shareis the emoji's count divided by all emoji tokens in the same period, which is what makes periods with different volumes comparable.top_nselects the emoji to follow, ranked over the whole corpus bymeasure, and the selected set is the same in every period.
Rows whose time is missing or unparseable contribute nothing. Glyphs are canonicalised through the package's codepoint key, so qualified and unqualified forms share one series.
See also
emoji_turnover() for vocabulary churn, emoji_seasonality() for
cyclical patterns.
Examples
df <- data.frame(
when = as.Date(c("2024-01-05", "2024-01-20", "2024-02-03")),
text = c("\U0001f600 hi", "\U0001f600\U0001f602", "\U0001f602 yes")
)
emoji_trend(df, text, when)
#> # A tibble: 4 × 5
#> .period emoji name n share
#> <date> <chr> <chr> <int> <dbl>
#> 1 2024-01-01 😀 grinning face 2 0.667
#> 2 2024-01-01 😂 face with tears of joy 1 0.333
#> 3 2024-02-01 😂 face with tears of joy 1 1
#> 4 2024-02-01 😀 grinning face 0 0