emoji_emotion_label() adds .emoji_emotion, the emotion with the highest
mean score among the row's emoji (using emoji_emotion()). Ties are broken
in Plutchik order; a row with nothing scorable, or with no emotion ahead of
the others, receives NA.
Arguments
- data
A data frame or tibble containing a text column. Grouped data frames are accepted. The verbs that work a row at a time (adding columns, or keeping and expanding rows) carry the grouping through to their result, as
dplyr::mutate()anddplyr::filter()do. The verbs that pool across rows – the counts, the co-occurrence edge lists, the time series – warn that they ignore the grouping and return one corpus-wide answer.- text
The text column to scan, supplied unquoted. Any atomic column is accepted and read as character, so a
factorworks and a numeric,Dateor logical one simply contains no emoji. A list column – or a data-frame column – is refused rather than coerced, because coercing one deparses it and the emoji found would be in the code rather than in your data. What counts as an emoji is the same in every verb; see the Detection section of tidyEmoji for the one case that surprises people, code points that are emoji only when they carryU+FE0F.- lexicon
Passed to
emoji_emotion().
Value
data, as a tibble, with .emoji_emotion (the winning emotion, or
NA when nothing was scorable) added, alongside the .emoji_n and
.emoji_n_scored counts it inherits from emoji_emotion(). The eight
per-emotion columns are not returned – the label is the point – unless
they were already in data, which is what
emoji_emotion() |> emoji_emotion_label() gives you: the profile and the
label side by side.
Details
Ties are broken in Plutchik order – the order the eight emotions are listed
in throughout the package (anger, anticipation, disgust, fear, joy, sadness,
surprise, trust) – so the winner is deterministic and does not depend on
the row's position in the data. It happens: 3 of the bundled lexicon's
150 glyphs tie for their top emotion, and because Plutchik order is
alphabetical the tie-break quietly favours the early names. U+1F3A4
scores anticipation and joy at 0.39 and is labelled anticipation; U+1F619
scores joy and trust at 0.83 and is labelled joy. So read
.emoji_n_scored alongside the label, and reach for emoji_emotion()
when a near-tie would change your reading: a single winning name cannot
show one.
A row whose scored emotions are all equal is the one case with no winner
to break a tie between, and it gets NA rather than the first name in the
order. An emoji scored zero on all eight is the obvious example. That
needs a custom lexicon to reach, the bundled one having no such glyph, and
.emoji_n_scored still separates it from a row with nothing to score.
See also
emoji_emotion() for the eight scores this collapses, and the
coverage caveat that applies to both; emoji_emotion_lexicon for the
underlying data; emoji_sentiment() for valence instead of emotion.
Examples
df <- data.frame(text = c("love it \U0001f60d", "scary \U0001f628", "meh"))
emoji_emotion_label(df, text)
#> # A tibble: 3 × 4
#> text .emoji_n .emoji_n_scored .emoji_emotion
#> <chr> <int> <int> <chr>
#> 1 love it 😍 1 1 joy
#> 2 scary 😨 1 1 fear
#> 3 meh 0 NA NA