has_cjk() reports, for each element, whether the string contains at least
one character from a CJK Unicode block.
Arguments
- x
A character vector. Anything else is coerced with
as.character(). That coercion is R's, not this package's, so a numeric vector is measured as R chooses to write it – which moves withoptions(scipen)andoptions(OutDec), and can therefore differ between sessions. Convert deliberately if you mean to measure numbers; these verbs are for text.
Details
"CJK" here means any block listed by cjk_blocks(), which includes CJK
punctuation and the halfwidth and fullwidth forms as well as the ideographs
and the phonetic scripts. That is deliberate – a column typed with a CJK
input method carries the punctuation too – but it does mean that a string
of nothing but ideographic full stops (U+3002) is TRUE. Use cjk_script()
when you need to know which kind of CJK you have.
See also
cjk_ratio() for how much of the text is CJK, cjk_script() for
which script it is.
Examples
# U+4E2D U+6587, "Chinese writing"
has_cjk(c("\u4e2d\u6587", "plain ASCII", NA))
#> [1] TRUE FALSE NA
# the ideographic full stop U+3002 counts, by design
has_cjk("\u3002")
#> [1] TRUE