Below we explore some of the capabilities of cppally, from its custom C++ scalar and vectors, to using templates and concepts.
Note: The classes r_vec and
r_vector are aliases of one another and thus can be used
interchangeably.
Registering C++ functions (to R)
To make a C++ function available to R we use the
[[cppally::register]] tag.
#include <cppally.hpp>
using namespace cppally;
[[cppally::register]]
void hello_world(){
print("Hello World!");
}After tagging our functions we want to make them available to R. To do that we have a few routes.
Registering C++ functions outside of a package context
After writing our hello world program in foo.cpp we can use
cpp_source() to compile and register the function to R.
cpp_source(file = "src/foo.cpp")Now the function is available in R
hello_world()
#> Hello World!Similarly we can use the helper cpp_eval to run simple
expressions and return the result without needing to include cppally.hpp
and register the function.
cpp_eval('print("Hello World Again!")')
#> Hello World Again!Note - For the rest of the examples it is assumed that the following code is always included beforehand.
Registering C++ functions inside a cppally-linked package
Since cppally is header-only, we can include the headers directly into our own package.
General steps to using cppally in a package
- Create package (if you haven’t already done so) using
usethis::create_tidy_package() - Run
cppally::use_cppally() - Run
cppally::document()
This will automatically add the necessary package content needed to
start working with cppally. For continuous development, use
cppally::load_all() to compile and register cppally tagged
functions, including our hello world function.
Note: We aim to integrate cppally registration into
the devtools framework for ease-of-use.
C++ types
cppally offers a rich set of R types in C++ that are NA-aware. This
means that common arithmetic and logical operations will account for
NA in a similar fashion to R.
Scalars
logical scalar - r_lgl
cppally’s scalar version of logical, r_lgl
can represent true, false or NA.
#> [1] TRUE
#> [1] FALSE
#> [1] NA
Logical operators work just like in R
[[cppally::register]]
r_vector<r_lgl> lgl_ops(){
return make_vec<r_lgl>(
r_true || r_false, // true
r_true && r_false, // false
r_na || r_true, // true
r_na && r_true, // NA
r_na && r_false, // false
r_na || r_na, // NA
r_na && r_na // NA
);
}
lgl_ops()
#> [1] TRUE FALSE TRUE NA FALSE NA NAUsing r_lgl in if-statements
For type-safety reasons r_lgl cannot be implicitly
converted to bool except in if-statements where an error is
thrown if the value is NA.
DON’T do this:
[[cppally::register]]
void bad_lgl_print(r_lgl condition){
if (condition){
print("true");
} else {
print("false");
}
}
bad_lgl_print(TRUE)
#> true
bad_lgl_print(FALSE)
#> false
bad_lgl_print(NA) # Can't implicitly convert NA to bool
#> Error:
#> ! Cannot implicitly convert r_lgl NA to bool, please checkDO this:
[[cppally::register]]
void good_lgl_print(r_lgl condition){
if (is_na(condition)){
print("NA");
} else if (condition){
print("true");
} else {
print("false");
}
}
good_lgl_print(TRUE)
#> true
good_lgl_print(FALSE)
#> false
good_lgl_print(NA) # NA is handled explicitly so no issues
#> NAWe can also use r_lgl members is_true() and
is_false() which return bool and are
equivalent to R’s isTRUE() and isFALSE()
[[cppally::register]]
void also_good_lgl_print(r_lgl condition){
if (condition.is_true()){
print("true");
} else {
print("not true");
}
}
also_good_lgl_print(TRUE)
#> true
also_good_lgl_print(FALSE)
#> not true
also_good_lgl_print(NA) # Falls into 'not true' branch here as expected
#> not trueImportant: The && and
|| operators for r_lgl do NOT
short-circuit like they do for bool. Both sides of
the expression are always evaluated. If you specifically require
short-circuiting behaviour, use is_true() and
is_false() before using && and
||.
All cppally scalar types are implemented as structs that contain the underlying C/C++ types as well as other member functions.
| cppally type | Description | Built on |
|---|---|---|
r_lgl |
Scalar logical | int |
r_int |
Scalar integer | int |
r_int64 |
Scalar 64-bit integer | int64_t |
r_dbl |
Scalar double | double |
r_str |
Scalar string | r_sexp |
r_str_view |
Scalar string (view) | SEXP |
r_cplx |
Scalar double complex | std::complex<double> |
r_raw |
Scalar raw | unsigned char |
r_sym |
Symbol | SEXP |
r_date |
Scalar date | r_dbl |
r_psxct |
Scalar date-time | r_dbl |
r_sexp |
Generic R object (SEXP)1 | SEXP |
C++ NA values and their R C API equivalents
| Type | Value | R C API Value |
|---|---|---|
r_lgl |
na<r_lgl>()/r_na
|
NA_LOGICAL |
r_int |
na<r_int>() |
NA_INTEGER |
r_int64 |
na<r_int64>() |
Not applicable |
r_dbl |
na<r_dbl>() |
NA_REAL |
r_str |
na<r_str>() |
NA_STRING |
r_cplx |
na<r_cplx>() |
Not applicable |
r_sym |
Not applicable | Not applicable |
r_sexp2 |
na<r_sexp>()/r_null
|
R_NilValue |
Checking equality
There are two ways to check for exact equality of cppally scalars -
with the == operator or with identical().
The cppally == operator always returns
r_lgl and identical() always returns
bool, which is a particularly important distinction when
dealing with NA values as the former can represent
NA while the latter cannot.
[[cppally::register]]
void cppally_equality(){
r_int x = na<r_int>();
r_int y = na<r_int>();
r_lgl x_equal_to_y = x == y;
bool x_identical_to_y = identical(x, y);
// NA so not printed
if ( x_equal_to_y.is_true() ){
print("x is equal to y\n");
}
// NA so not printed
if ( x_equal_to_y.is_false() ){
print("x is not equal to y\n");
}
// NA so printed
if (is_na(x_equal_to_y)){
print("`x == y` produces `NA`\n");
}
// Both na<r_int>() therefore they are identical to each other
if (x_identical_to_y){
print("x is identical to y\n");
}
}
cppally_equality()
#> `x == y` produces `NA`
#> x is identical to yidentical() can not only compare scalars, but also
vectors, lists, factors, and data frames.
Scalar operators
cppally also defines arithmetic and relational comparison operators
for its scalar types. Like the logical and equality operators seen
earlier, they are all NA-aware.
Scalar arithmetic operators
Addition
#> [1] 2.5
Subtraction
#> [1] -1
Multiplication
#> [1] 6
Division
#> [1] 3
Addition with NA
#> [1] NA
Subtraction with NA
#> [1] NA
Multiplication with NA
#> [1] NA
Division with NA
#> [1] NA
Scalar relational operators
Less than
#> [1] TRUE
Less than or equal to
#> [1] TRUE
Greater than
#> [1] TRUE
Greater than or equal to
#> [1] FALSE
Less than with NA
#> [1] NA
Less than or equal to with NA
#> [1] NA
Greater than with NA
#> [1] NA
Greater than or equal to with NA
#> [1] NA
Other defined operators not showcased: ++,
--, +=, -=, *=,
/=, %=, -, %,
|, &, !
Vectors
cppally vectors are templated and can be thought of as containers of
scalar elements like r_int, r_dbl, etc.
We can create vectors like so
// Integer vector of size n
[[cppally::register]]
r_vector<r_int> new_integer_vector(int n){
r_vector<r_int> int_vctr(n, /*fill = */ r_int(0));
return int_vctr;
}
new_integer_vector(3)
#> [1] 0 0 0inline vectors
To create inline vectors, use make_vec<>
#> [1] 1.0 1.5 2.0 NA
We can add names on the fly with arg()
make_vec<r_dbl>(
arg("first") = 1,
arg("second") = 1.5,
arg("third") = 2,
arg("last") = na<r_dbl>()
)#> first second third last
#> 1.0 1.5 2.0 NA
In R a list is a generic vector, so cppally defines lists as
r_vector<r_sexp>, a vector of the generic type
r_sexp.
#> [[1]]
#> [1] 1
#>
#> [[2]]
#> [1] 2
#>
#> [[3]]
#> [1] 3
A list of all cppally vectors of length 0
[[cppally::register]]
r_vector<r_sexp> all_vectors(){
return make_vec<r_sexp>(
arg("logical") = r_vector<r_lgl>(),
arg("integer") = r_vector<r_int>(),
arg("integer64") = r_vector<r_int64>(), // Requires bit64
arg("double") = r_vector<r_dbl>(),
arg("character") = r_vector<r_str>(),
arg("character") = r_vector<r_str_view>(),
arg("raw") = r_vector<r_raw>(),
arg("date") = r_vector<r_date>(),
arg("date-time") = r_vector<r_psxct>(),
arg("list") = r_vector<r_sexp>()
);
}
all_vectors()
#> $logical
#> logical(0)
#>
#> $integer
#> integer(0)
#>
#> $integer64
#> integer64(0)
#>
#> $double
#> numeric(0)
#>
#> $character
#> character(0)
#>
#> $character
#> character(0)
#>
#> $raw
#> raw(0)
#>
#> $date
#> Date of length 0
#>
#> $`date-time`
#> POSIXct of length 0
#>
#> $list
#> list()Scalar math
There is a rich suite of math functions that accept cppally types.
[[cppally::register]]
r_vector<r_dbl> cppally_math(r_dbl x){
return make_vec<r_dbl>(
arg("abs") = abs(x),
arg("floor") = floor(x),
arg("ceiling") = ceiling(x),
arg("trunc") = trunc(x),
arg("round") = round(x),
arg("signif") = signif(x, 3),
arg("sign") = sign(x),
arg("min") = min(0, x),
arg("max") = max(0, x),
arg("sqrt") = sqrt(x),
arg("pow") = pow(x, 2),
arg("exp") = exp(x),
arg("log") = log(x),
arg("log_base") = log(x, 2),
arg("log10") = log10(x)
);
}
cppally_math(2.5)
#> abs floor ceiling trunc round signif sign
#> 2.5000000 2.0000000 3.0000000 2.0000000 2.0000000 2.5000000 1.0000000
#> min max sqrt pow exp log log_base
#> 0.0000000 2.5000000 1.5811388 6.2500000 12.1824940 0.9162907 1.3219281
#> log10
#> 0.3979400
cppally_math(NA)
#> abs floor ceiling trunc round signif sign min
#> NA NA NA NA NA NA NA NA
#> max sqrt pow exp log log_base log10
#> NA NA NA NA NA NA NACoercion
To coerce from one scalar to another we can use
as<T>
double_to_int(pi)
#> [1] 3
double_to_int(NA_real_)
#> [1] NAWe can also coerce from one vector type to another
[[cppally::register]]
r_vector<r_int> to_int_vec(r_vector<r_dbl> x){
return as<r_vector<r_int>>(x);
}
to_int_vec(c(0, 1.5, NA))
#> [1] 0 1 NASince as<T> is extremely flexible, we can also
coerce from a scalar to a vector or vice versa
[[cppally::register]]
r_vector<r_sexp> coercions(){
r_dbl a(4.2);
r_vector<r_dbl> b = make_vec<r_dbl>(2.5);
return make_vec<r_sexp>(
as<r_vector<r_int>>(a),
as<r_int>(a),
as<r_int>(b),
as<r_dbl>(b)
);
}
coercions()
#> [[1]]
#> [1] 4
#>
#> [[2]]
#> [1] 4
#>
#> [[3]]
#> [1] 2
#>
#> [[4]]
#> [1] 2.5We can even coerce to and from C++ vectors
[[cppally::register]]
r_vector<r_int> cpp_vectors_example(r_vector<r_int> x){
std::vector x_cpp = as<std::vector<r_int>>(x);
x_cpp.push_back(r_int(42));
return as<r_vector<r_int>>(x_cpp);
}
cpp_vectors_example(41L)
#> [1] 41 42While coercing to a std::vector just to push back an
element before coercing back might not be the most efficient, it does
showcase how easy it is to work with cppally vectors and C++
vectors.
Strings
cppally provides the useful string type r_str
We can create R strings easily
#> [1] "hello"
To get a C or C++ string, use the members c_str() and
cpp_str() respectively
C string via c_str()
#> [1] "hello"
C++ string_view via cpp_str()
This can be converted into a std::string via its constructor
[[cppally::register]]
r_str str_concatenate(r_str x, r_str y, r_str sep){
std::string left = std::string(x.cpp_str());
std::string right = std::string(y.cpp_str());
std::string middle = std::string(sep.cpp_str());
std::string combined = left + middle + right;
return r_str(combined.c_str());
}
str_concatenate("hello", "how are you?", sep = ", ")
#> [1] "hello, how are you?"Dates and date-times
cppally offers two classes: r_date for dates and
r_psxct for date-times (based on POSIX time). These classes
offer a rich set of member functions to perform date arithmetic and
manipulation.
To quickly re-familiarise ourselves - R Dates and Date-times are
essentially double-valued offsets from Unix time. Unix time is defined
as the number of seconds since 1 Jan 1970, and so r_date is
effectively an offset representing the number of days since epoch, and
similarly r_psxct is an offset representing the number of
seconds since epoch.
To create a new r_date, use the
r_date(year, month, day) constructor.
[[cppally::register]]
r_date create_date(int year, int month, int day){
return r_date(year, month, day);
}
[[cppally::register]]
r_date date_from_epoch_offset(double n){
return r_date(n);
}
[[cppally::register]]
r_psxct create_datetime(int year, int month, int day, int hour, int minute, double second){
return r_psxct(year, month, day, hour, minute, second);
}
[[cppally::register]]
r_psxct datetime_from_epoch_offset(double n){
return r_psxct(n);
}
create_date(1970, 1, 1); # This builds the Unix epoch date
#> [1] "1970-01-01"You can also build an r_date from an epoch offset.
date_from_epoch_offset(0) # Also builds the Unix epoch
#> [1] "1970-01-01"Similarly, to create a new r_psxct, use the
r_psxct(year, month, day, hour, minute, second)
constructor.
create_datetime(1970, 1, 1, 0, 0, 0); # Unix epoch date-time
#> [1] "1970-01-01 UTC"
datetime_from_epoch_offset(0) # Unix epoch from an offset
#> [1] "1970-01-01 UTC"To get the current date and date-time, use
r_date::today() and r_psxct::now()
respectively.
[[cppally::register]]
r_date get_today(){
return r_date::today();
}
[[cppally::register]]
r_psxct get_now(){
return r_psxct::now();
}
get_today()
#> [1] "2026-09-10"
get_now()
#> [1] "2026-09-10 17:38:38 UTC"Note: r_psxct currently only supports
UTC and no other time-zones.
Adding days, months, and years
r_date and r_psxct both have the template
function add<>, which allows us to add years, months,
weeks, days, hours, minutes and seconds to our dates and date-times.
[[cppally::register]]
r_date add_days(r_date x, double n){
return x.add<"days">(n);
}
[[cppally::register]]
r_date add_months(r_date x, double n){
return x.add<"months">(n);
}
y2k <- create_date(2000, 1, 1)
# Day after y2k
y2k |>
add_days(1)
#> [1] "2000-01-02"
# Week after y2k
y2k |>
add_days(7)
#> [1] "2000-01-08"
# Month after y2k
y2k |>
add_months(1)
#> [1] "2000-02-01"
# Year after y2k
y2k |>
add_months(12)
#> [1] "2001-01-01"Accessing date components
Those familiar with lubridate will instantly be familiar with the naming of these functions, something that was an intentional design choice to reduce confusion for those who are familiar with tidyverse.
[[cppally::register]]
r_int get_year(r_date x){
return x.year();
}
[[cppally::register]]
r_int get_month(r_date x){
return x.month();
}
[[cppally::register]]
r_int get_week(r_date x){
return x.week();
}
[[cppally::register]]
r_int get_day(r_date x){
return x.day();
}
[[cppally::register]]
r_int get_wday(r_date x, int week_start){
return x.wday(week_start);
}
[[cppally::register]]
r_int get_yday(r_date x){
return x.yday();
}
[[cppally::register]]
r_int get_iso_week(r_date x){
return x.iso_week();
}
[[cppally::register]]
r_int get_iso_year(r_date x){
return x.iso_year();
}
get_year(y2k)
#> [1] 2000
get_month(y2k)
#> [1] 1
get_week(y2k)
#> [1] 1
get_day(y2k) # Day of the month
#> [1] 1
get_wday(y2k, week_start = 1) # week_start = 1 = Mon
#> [1] 6
get_yday(y2k)
#> [1] 1We can even retrieve the ISO week and ISO year.
get_iso_week(y2k) # Last (52nd) ISO week of the ISO year
#> [1] 52
get_iso_year(y2k) # ISO year is still 1999
#> [1] 1999Date rounding
We can also round dates using the member functions
floor(), ceiling(), and
round().
[[cppally::register]]
r_date floor_month(r_date x){
return x.floor<"months">();
}
[[cppally::register]]
r_date ceiling_year(r_date x){
return x.ceiling<"years">();
}
[[cppally::register]]
r_date round_week(r_date x, int week_start){
return x.round<"weeks">(week_start);
}
# Start of the month after advancing 100 days
y2k |>
add_days(100) |>
floor_month()
#> [1] "2000-04-01"
# ceiling() does not switch on boundary
y2k |>
ceiling_year()
#> [1] "2000-01-01"
y2k |>
add_days(1) |>
ceiling_year()
#> [1] "2001-01-01"
y2k |>
round_week(week_start = 1)
#> [1] "2000-01-03"Difference between two dates
To calculate the time difference between two dates or date-times, use
time_diff<>.
[[cppally::register]]
r_dbl diff_days(r_date x, r_date y){
return time_diff<"days">(x, y);
}
[[cppally::register]]
r_dbl diff_months(r_date x, r_date y){
return time_diff<"months">(x, y);
}
x <- y2k
y <- x |> add_months(3)
diff_days(x, y)
#> [1] 91
diff_months(x, y)
#> [1] 3Fractional months
Both add<"months"> and
time_diff<"months"> can deal with fractional
months.
The previous sections dealt with dates, so we need to define
functions that work with r_psxct (date-times).
[[cppally::register]]
r_psxct as_dt(r_date x){
return x.as_datetime();
}
[[cppally::register]]
r_psxct dt_add_months(r_psxct x, double n){
return x.add<"months">(n);
}
[[cppally::register]]
r_dbl dt_diff_months(r_psxct x, r_psxct y){
return time_diff<"months">(x, y);
}
y2k_datetime <- as_dt(y2k)
mid_month <- y2k_datetime |>
dt_add_months(0.5)
# Half-way through the month
mid_month
#> [1] "2000-01-16 12:00:00 UTC"The difference between the start of the month and halfway through the month should return 0.5, confirming that fractional month arithmetic is working.
dt_diff_months(y2k_datetime, mid_month)
#> [1] 0.5Symbols
Symbols have class r_sym and can be created directly
from a string literal
#> new_symbol
Or from a cppally string
#> symbol_from_string
Cached strings & symbols
cppally provides an efficient caching strategy for constructing cppally strings/symbols from string literals
cached_str<>
#> [1] "cached_string"
This initialises the string once, caches it (to R’s CHARSXP pool), and efficiently re-uses the cached string for each subsequent call.
We can cache symbols in a similar way
#> cached_symbol
Lists
r_sexp is generally interpreted as an “element of a
list” since lists are defined as r_vector<r_sexp>, a
vector that holds generic r_sexp elements.
new_list(0)
#> list()
new_list(3)
#> [[1]]
#> NULL
#>
#> [[2]]
#> NULL
#>
#> [[3]]
#> NULLThe problem with a class like r_sexp is that it is by
design generic and therefore difficult to work with in C++. To
disambiguate the actual type we can use r_sexp_visit() via
a C++ lambda.
Example: using r_sexp_visit() to resize
every vector to length n in-place
Factors
We can create a factor via r_factors()
new_factor(letters)
#> [1] a b c d e f g h i j k l m n o p q r s t u v w x y z
#> Levels: a b c d e f g h i j k l m n o p q r s t u v w x y zIn cppally, like R, factors are not vectors and therefore do not
satisfy the RVector concept. To access the underlying integer codes
vector, use the public codes() member function
static_assert(!RVector<r_factors>);
[[cppally::register]]
r_vector<r_int> factor_codes(r_factors x){
return x.codes();
}
letter_fct <- new_factor(letters)
letter_fct |>
factor_codes()
#> [1] 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25
#> [26] 26Calling R functions
cppally provides the r_function class, which can be used
to easily call any R function.
[[cppally::register]]
r_date r_get_today(){
r_function sys_date("Sys.Date");
return as<r_date>(sys_date());
}
r_get_today()
#> [1] "2026-09-10"To get a function from a specific package, use pkg_env,
a helper that returns the namespace of a package. The result of
pkg_env can be passed to the env (or second)
arg of r_function’s symbol and string constructors.
[[cppally::register]]
r_function base_sum_fn(){
return r_function(cached_sym<"sum">(), pkg_env<"base">());
}
base_sum_fn()
#> function (..., na.rm = FALSE) .Primitive("sum")r_function also allows named arguments to be passed.
[[cppally::register]]
r_dbl base_sum(r_vector<r_dbl> x, bool na_rm){
return as<r_dbl>(base_sum_fn()(x, arg("na.rm") = na_rm));
}
base_sum(c(NA, 1:3, NA), na_rm = TRUE)
#> [1] 6Value matching
Use r_vector member find() to find the
0-indexed locations of a scalar value.
[[cppally::register]]
r_vector<r_int> find_empty_string(r_vector<r_str> x){
return x.find(r_str(""));
}
x <- c("zero", "one", "two", "three", "four")
find_empty_string(x)
#> integer(0)
# Add empty strings
x[c(1, 3)] <- ""
find_empty_string(x)
#> [1] 0 2To find the locations of multiple values (first match), use
match(). It works like R’s base::match() but
is 0-indexed.
[[cppally::register]]
r_vector<r_int> match_strs(r_vector<r_str> x, r_vector<r_str> table){
return match(x, table);
}
letters
#> [1] "a" "b" "c" "d" "e" "f" "g" "h" "i" "j" "k" "l" "m" "n" "o" "p" "q" "r" "s"
#> [20] "t" "u" "v" "w" "x" "y" "z"
vowels <- c("a", "e", "i", "o", "u")
# cppally::match is 0-indexed
match_strs(vowels, letters)
#> [1] 0 4 8 14 20cppally provides the IS_IN infix operator, identical to
R’s %in%
[[cppally::register]]
r_vector<r_lgl> cpp_in(r_vector<r_str> x, r_vector<r_str> table){
return x IS_IN table;
}
cpp_in(c("a", "A", NA), letters)
#> [1] TRUE FALSE FALSETo mimic R’s new %notin% operator, simply use the
logical negation operator.
[[cppally::register]]
r_vector<r_lgl> cpp_not_in(r_vector<r_str> x, r_vector<r_str> table){
return !(x IS_IN table);
}Technical note: cppally internally negates the
intermediate result of x IS_IN table in-place in this
particular case because it satisfies specific properties of exclusivity
which while not covered here, may be covered in a later vignette. This
in-place negation is naturally efficient as it avoids allocating a new
vector.
cpp_not_in(c("a", "A", NA), letters)
#> [1] FALSE TRUE TRUESubsetting
Subsetting is 0-indexed in cppally, as is all other indexing.
template <RVector T, typename subscript_t>
requires any<subscript_t, r_lgl, r_int, r_str>
[[cppally::register]]
T cpp_subset(T x, r_vector<subscript_t> y){
return subset(x, y);
}Subsetting with integer indices
x <- 1:10
cpp_subset(x, 0L) # index 0 is 1st value
#> [1] 1
cpp_subset(x, 9L) # index 9 is last value here
#> [1] 10
cpp_subset(x, 10L) # index 10 is out-of-bounds so NA is returned
#> [1] NA
cpp_subset(x, NA_integer_) # NA is returned with integer NA index
#> [1] NASubsetting with logical indices
cpp_subset(x, x > 5)
#> [1] 6 7 8 9 10
cpp_subset(x, x > 100)
#> integer(0)
# It differs to base subsetting (via `[`)
cpp_subset(x, rep(NA, 10)) # cppally only returns values associated with TRUE
#> integer(0)
x[rep(NA, 10)] # base R returns NA values when subsetting with NA logicals
#> [1] NA NA NA NA NA NA NA NA NA NAbase R performs negative subsetting by using negative numbers. This
is not possible with our 0-indexed subset as we would have to also
represent negative zero to subset everything except the first element at
index 0. Instead cppally has an invert argument.
template <RVector T, typename subscript_t>
requires any<subscript_t, r_lgl, r_int, r_str>
[[cppally::register]]
T cpp_negative_subset(T x, r_vector<subscript_t> y){
return subset(x, y, /*invert=*/ true);
}
cpp_negative_subset(x, 0L) # Everything but 1st
#> [1] 2 3 4 5 6 7 8 9 10
cpp_negative_subset(x, x > 5) # Everything except where x > 5
#> [1] 1 2 3 4 5
cpp_negative_subset(x, integer()) # Everything
#> [1] 1 2 3 4 5 6 7 8 9 10Named subsetting is also supported. cppally internally hashes the vector names and performs hash lookups. For more info, see the ‘Vector Names Hashing’ vignette.
Concepts and Templates
One of the most powerful features of C++20 are concepts. These allow users to write human-readable templates and constraints.
When writing your own templates, it is necessary to place them in headers for cppally registration to work correctly.
Let’s practice by creating the abs() function in C++
using templates and the RNumber concept.
template <RNumber T>
[[cppally::register]]
T cpp_abs(T x){
if (is_na(x)){
return na<T>();
} else if (x < 0){
return -x;
} else {
return x;
}
}What’s nice is that it works correctly for integers and doubles while simultaneously preserving their type
cpp_abs(-4.2)
#> [1] 4.2
cpp_abs(-3L)
#> [1] 3
class(cpp_abs(-4.2)) # Double preserved
#> [1] "numeric"
class(cpp_abs(-3L)) # Integer preserved
#> [1] "integer"This type of programming is historically tricky within the R C API
and typically necessitates a switch statement that switches on the
object’s type, handling each type separately. With our
abs() template, the logic is correctly handled with one set
of operations.
How it works
The top-line template <RNumber T> declares a
template that encapsulates T, an RNumber - a
concept that contains r_int, r_int64 and
r_dbl
If x is NA then we immediately also return NA via
na<T>() which is a templated function that returns NA
of the input type T.
To correctly register templates, the ‘[[cppally::register]]’ tag must always go above the function name.
Templates without function arguments
Explicit instantiation (from R) is unfortunately not possible and template types must be deduced from supplied arguments.
Here foo() will not be compiled because the function has
no arguments that let the compiler automatically deduce what
T is. In C++ you would always call this function like so:
foo<T>(). Unfortunately we can’t do that from R
directly.
You may get a cryptic compiler error like this
along with an equally cryptic note
Even though these kinds of templates can be written with cppally in C++, they cannot be exported to R.
An obvious and somewhat ugly workaround is to include a prototype argument that allows the template parameter to be deduced from.
// Return the default constructor result of RScalar types
template <RScalar T>
[[cppally::register]]
T scalar_default(T ptype){
return T();
}
scalar_default(integer(1)) # Default is 0L
#> [1] 0
scalar_default(numeric(1)) # Default is 0.0
#> [1] 0
scalar_default(character(1)) # Default is ""
#> [1] ""Exporting variadic templates are also not supported. The best
alternative is to use lists (r_vector<r_sexp>).
In the above example we used the RScalar concept which
includes all cppally scalar types (excluding r_sexp). For a
list of all cppally concepts, please see the Annex
Attributes
Attributes can be manipulated via functions defined in the attr namespace.
Example: Adding names to a list
[[cppally::register]]
r_vector<r_sexp> set_list_names(r_vector<r_sexp> x, r_vector<r_str> names){
x.set_names(names);
return x;
}
set.seed(42)
norm_samples <- lapply(1:5, \(x) rnorm(10, mean = x))
set_list_names(norm_samples, paste0("sample_", 1:5))
#> $sample_1
#> [1] 2.3709584 0.4353018 1.3631284 1.6328626 1.4042683 0.8938755 2.5115220
#> [8] 0.9053410 3.0184237 0.9372859
#>
#> $sample_2
#> [1] 3.3048697 4.2866454 0.6111393 1.7212112 1.8666787 2.6359504
#> [7] 1.7157471 -0.6564554 -0.4404669 3.3201133
#>
#> $sample_3
#> [1] 2.693361 1.218692 2.828083 4.214675 4.895193 2.569531 2.742731 1.236837
#> [9] 3.460097 2.360005
#>
#> $sample_4
#> [1] 4.455450 4.704837 5.035104 3.391074 4.504955 2.282991 3.215541 3.149092
#> [9] 1.585792 4.036123
#>
#> $sample_5
#> [1] 5.205999 4.638943 5.758163 4.273295 3.631719 5.432818 4.188607 6.444101
#> [9] 4.568554 5.655648More useful attribute helpers
-
get_attrs()- Returns a list of attributes (possiblyr_vector<r_sexp>(r_null)) -
set_attrs()- Sets attributes to ones specified. Note: replaces any current attributes -
clear_attrs()- Removes all attributes -
set_attr()- Set a single attribute -
get_attr()- Get a single attribute -
inherits1()- Does object inherit class? -
inherits_any()- Does object inherit at least one of the specified classes? -
inherits_all()- Does object inherit all of the specified classes? -
modify_attrs()- Modifies current attributes but doesn’t remove any existing ones
Regular sequences
There are two core functions for generating regular sequences -
sequence() and seq(). seq()
behaves exactly like the R equivalent base::seq(), and
sequence() behaves like the R equivalent
base::sequence(), with the exception that it accepts scalar
arguments instead of vector ones, and also works with decimal
increments.
#> [1] 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0
#> [1] 0 -1 -2 -3 -4
We can also use sequence() to easily replicate
base::seq_len()
[[cppally::register]]
r_vector<r_int> cpp_seq_len(r_size_t n){
return sequence(n, /* from = */ r_int(1), /* by = */ r_int(1));
}
cpp_seq_len(5)
#> [1] 1 2 3 4 5It is also straightforward to replicate base::sequence()
with pmap()
Using multiple OpenMP threads
To set the number of global OpenMP threads, use
cppally::set_threads().
To see the number of currently available threads, use
cppally::get_threads().
From R we can then register these C++ function and use them.
[[cppally::register]]
void cpp_set_threads(int n){
set_threads(n);
}
[[cppally::register]]
int cpp_get_threads(){
return get_threads();
}
cpp_set_threads(2) # Set to 2 threads
cpp_get_threads() # Number of threads now available
#> [1] 2
cpp_set_threads(1) # Reset back to 1
cpp_get_threads() # Should return 1
#> [1] 1Because this is unique to each dll file, setting threads in one R package doesn’t affect another. Once threads have been set, all cppally code that can make use of them, will use them. This means you should speed improvements across many cppally functions.
Stats functions
cppally provides some stats functions which are both common and difficult to implement efficiently and safely.
-
sum()- Sum of values -
mean()- Average of values -
range()- Min and max range of values -
var()- Variance
template <RNumber T>
[[cppally::register]]
r_dbl cpp_sum(r_vector<T> x, bool na_rm){
return as<r_dbl>(sum(x, na_rm));
}
template <RNumber T>
[[cppally::register]]
r_vector<T> cpp_range(r_vector<T> x, bool na_rm){
return range(x, na_rm);
}
template <RNumber T>
[[cppally::register]]
r_dbl cpp_var(r_vector<T> x, bool na_rm){
return var(x, na_rm);
}
library(bench)
set.seed(1)
x <- sample.int(10, 5e05, replace = TRUE)
x[sample.int(10^5, 10^4)] <- NA # Add many NAs
# Sum (ignoring NAs)
mark(
cppally_sum = cpp_sum(x, na_rm = TRUE),
base_sum = sum(x, na.rm = TRUE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_sum 109µs 109µs 9007. 0B 0
#> 2 base_sum 313µs 315µs 3141. 0B 0
# Sum (not ignoring NAs)
mark(
cppally_sum = cpp_sum(x, na_rm = FALSE),
base_sum = sum(x, na.rm = FALSE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_sum 861ns 1.05µs 885265. 0B 0
#> 2 base_sum 160ns 181.03ns 4330836. 0B 0
# Range (ignoring NAs)
mark(
cppally_range = cpp_range(x, na_rm = TRUE),
base_range = range(x, na.rm = TRUE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_range 126.36µs 126.9µs 7417. 0B 0
#> 2 base_range 2.36ms 2.4ms 392. 11.4MB 258.
# Range (not ignoring NAs)
mark(
cppally_range = cpp_range(x, na_rm = FALSE),
base_range = range(x, na.rm = FALSE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_range 861ns 931.09ns 1005228. 0B 0
#> 2 base_range 442µs 1.04ms 1056. 1.91MB 57.7
# Variance (ignoring NAs)
mark(
cppally_var = cpp_var(x, na_rm = TRUE),
base_var = var(x, na.rm = TRUE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_var 591.63µs 649.2µs 1539. 0B 0
#> 2 base_var 4.44ms 5.1ms 199. 5.74MB 27.4
# Variance (not ignoring NAs)
mark(
cppally_var = cpp_var(x, na_rm = FALSE),
base_var = var(x, na.rm = FALSE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_var 881ns 961.01ns 975679. 0B 0
#> 2 base_var 988µs 1.07ms 928. 3.81MB 84.4
# Using multiple threads
cpp_set_threads(4)
# Sum (ignoring NAs)
mark(
cppally_sum = cpp_sum(x, na_rm = TRUE),
base_sum = sum(x, na.rm = TRUE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_sum 57.5µs 58.4µs 16456. 0B 0
#> 2 base_sum 313.4µs 314.6µs 3133. 0B 0
# Sum (not ignoring NAs)
mark(
cppally_sum = cpp_sum(x, na_rm = FALSE),
base_sum = sum(x, na.rm = FALSE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_sum 861ns 932ns 937173. 0B 0
#> 2 base_sum 160ns 181ns 2872149. 0B 0
# Range (ignoring NAs)
mark(
cppally_range = cpp_range(x, na_rm = TRUE),
base_range = range(x, na.rm = TRUE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_range 168.39µs 169.8µs 5767. 0B 0
#> 2 base_range 2.37ms 2.4ms 397. 11.4MB 271.
# Range (not ignoring NAs)
mark(
cppally_range = cpp_range(x, na_rm = FALSE),
base_range = range(x, na.rm = FALSE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_range 861ns 931ns 999665. 0B 0
#> 2 base_range 431µs 450µs 2189. 1.91MB 96.1
# Variance (ignoring NAs)
mark(
cppally_var = cpp_var(x, na_rm = TRUE),
base_var = var(x, na.rm = TRUE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_var 307.95µs 324.33µs 3045. 0B 0
#> 2 base_var 3.72ms 3.76ms 265. 5.72MB 43.8
# Variance (not ignoring NAs)
mark(
cppally_var = cpp_var(x, na_rm = FALSE),
base_var = var(x, na.rm = FALSE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_var 871ns 961ns 979359. 0B 0
#> 2 base_var 352µs 372µs 2656. 3.81MB 243.
# Reset to single-threaded
cpp_set_threads(1)A note on cppally::sum
cppally::sum() is written to use SIMD instructions for
performance where possible. It also uses higher width integers for
accuracy when dealing with integer vectors. For example, if summing a
32-bit integer vector, it uses intermediate 64-bit integers, and
likewise if summing a 64-bit integer vector, it uses intermediate
128-bit integers (if your compiler supports it). Doing it this way
avoids integer overflow entirely for vectors whose length is less than
2^32. It also avoids intermediate integer overflow that would have
arisen if same-width integers were used. For example, for
x <- .Machine$integer.max * c(-1L, -1L, 1L, 1L), the sum
of x is trivially 0, but the intermediate addition of x[1]
and x[2] would result in overflow if one were to use 32-bit
integers.
Unique values
n_unique() is a convenient helper to calculate the
number of unique values. unique() is also highly
performant.
template <RVector T>
[[cppally::register]]
r_int cpp_n_unique(T x){
return as<r_int>(n_unique(x));
}
template <RVector T>
[[cppally::register]]
T cpp_unique(T x, bool sort){
return unique(x, sort);
}
x <- sample(1:100, 10^5, replace = TRUE)
mark(
base_n_unique = length(unique(x)),
cppally_n_unique = cpp_n_unique(x)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 base_n_unique 580µs 592µs 1688. 1.38MB 54.4
#> 2 cppally_n_unique 112µs 112µs 8762. 0B 0
mark(
base_unique = unique(x),
cppally_unique = cpp_unique(x, sort = FALSE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 base_unique 552µs 567µs 1708. 1.38MB 52.2
#> 2 cppally_unique 112µs 113µs 8732. 448B 0
mark(
base_sorted_unique = sort(unique(x)),
cppally_sorted_unique = cpp_unique(x, sort = TRUE)
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 base_sorted_unique 580µs 593µs 1681. 1.38MB 49.6
#> 2 cppally_sorted_unique 113µs 114µs 8633. 896B 0Sorting
cppally also offers order() and sort() for
sorting purposes. A hybrid approach is used, incorporating various
sorting algorithms, including ska sort (a variant on radix sort),
counting sort, quick sort, etc. These have been implemented with speed
being a key focus.
template <RSortableVector T>
[[cppally::register]]
r_vector<r_int> cpp_order(T x){
// Add 1 to return 1-indexed permutation to match base::order()
// In general C++ code, always use 0-indexed as cppally indexing is
// always 0-indexed.
return order(x) + 1;
}
template <RSortableVector T>
[[cppally::register]]
T cpp_sort(T x){
return sort(x);
}cppally::sort always sorts NA values last
and so is equivalent to
base::sort(..., method = "radix", na.last = TRUE).
r_radix_sort <- function(x){
sort(x, method = "radix", na.last = TRUE)
}
r_radix_order <- function(x){
order(x, method = "radix", na.last = TRUE)
}
# Small data
mark(cpp_sort(c(3, 2, 1)), r_radix_sort(c(3, 2, 1)))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_sort(c(3, 2, 1)) 961.1ns 1.05µs 794045. 0B 0
#> 2 r_radix_sort(c(3, 2, 1)) 19.7µs 22.44µs 43219. 0B 17.3
mark(cpp_order(c(3, 2, 1)), r_radix_order(c(3, 2, 1)))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_order(c(3, 2, 1)) 961.01ns 1.12µs 788512. 0B 0
#> 2 r_radix_order(c(3, 2, 1)) 9.06µs 10.36µs 92789. 0B 18.6
# Integer data
ints <- sample.int(20, 2e05, replace = TRUE)
mark(cpp_sort(ints), r_radix_sort(ints))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_sort(ints) 456µs 984.5µs 1292. 1.53MB 43.8
#> 2 r_radix_sort(ints) 860µs 1.43ms 756. 1.53MB 25.5
mark(cpp_order(ints), r_radix_order(ints))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_order(ints) 396µs 411µs 2411. 781KB 40.9
#> 2 r_radix_order(ints) 490µs 511µs 1951. 781KB 35.8
# Double-floating point precision data
dbls <- as.double(ints)
mark(cpp_sort(dbls), r_radix_sort(dbls))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_sort(dbls) 1.38ms 1.41ms 711. 2.29MB 48.4
#> 2 r_radix_sort(dbls) 3.75ms 3.96ms 252. 2.29MB 19.0
mark(cpp_order(dbls), r_radix_order(dbls))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_order(dbls) 1.3ms 1.32ms 758. 781KB 14.9
#> 2 r_radix_order(dbls) 3.15ms 3.42ms 294. 781KB 6.22
# Floating point data with decimals
dcmls <- rnorm(length(dbls))
mark(cpp_sort(dcmls), r_radix_sort(dcmls))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_sort(dcmls) 4.63ms 4.7ms 213. 2.29MB 15.5
#> 2 r_radix_sort(dcmls) 6.04ms 6.2ms 160. 2.29MB 8.64
mark(cpp_order(dcmls), r_radix_order(dcmls))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_order(dcmls) 4.6ms 4.62ms 215. 781KB 4.10
#> 2 r_radix_order(dcmls) 5.23ms 5.7ms 178. 781KB 4.14
# String data
strs <- round(dcmls, 2) |> paste0()
mark(cpp_sort(strs), r_radix_sort(strs))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_sort(strs) 1.81ms 2.26ms 441. 2.29MB 32.9
#> 2 r_radix_sort(strs) 1.82ms 2.19ms 465. 2.29MB 41.4
mark(cpp_order(strs), r_radix_order(strs))
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cpp_order(strs) 1.03ms 1.1ms 836. 781KB 22.3
#> 2 r_radix_order(strs) 772.68µs 1.08ms 949. 781KB 25.0Grouping
Some areas of analysis require working with groups, with sometimes
very large numbers of them. cppally provides a powerful class
groups, along with make_groups(), a function
that creates groups from vectors.
The structure of groups is:
-
r_vector<r_int> ids- The group IDs -
int n_groups- Number of unique groups -
bool ordered- Do the group IDs specify a sorting order, or are they by order-of-first-appearance? -
bool sorted- Are the group IDs sorted? (This can also be true for order-of-first-appearance IDs) -
r_vector<r_int> starts()- Returns a vector of start locations of each unique group, signifying the location in the data at which each group initially appeared -
r_vector<r_int> counts()- Returns a vector of frequency counts of each unique group -
r_vector<r_int> order()- Returns a 0-indexed permutation vector that can be used to sort group IDs
Note: When the groups are in sorted-order (where
sorted is true), member functions like
starts() and counts() use a hybrid search
strategy combining linear, gallop, and binary searching. This strategy
ensures that search time complexity scales with the number of groups and
not the size of the data.
template <RVector T>
[[cppally::register]]
r_vector<r_sexp> group_info(T x){
groups g = make_groups(x, /*ordered=*/ false);
return make_vec<r_sexp>(
arg("group_id") = g.ids,
arg("n_groups") = g.n_groups,
arg("ordered") = g.ordered,
arg("sorted") = g.sorted,
arg("starts") = g.starts(),
arg("counts") = g.counts(),
arg("order") = g.order(),
arg("group_names") = group_names(x, g)
);
}
set.seed(1)
x <- sample(c("A", "B", "C"), 20, replace = TRUE)
# Convert into group info
letter_groups <- group_info(x)
letter_groups$group_id # Group IDs
#> [1] 0 1 0 2 0 1 1 2 2 1 1 0 0 0 2 2 2 2 1 0
letter_groups$n_groups # Number of groups
#> [1] 3
letter_groups$ordered # Groups were collected using order-of-first-appearance
#> [1] FALSE
letter_groups$sorted # Are group IDs in sorted order?
#> [1] FALSE
letter_groups$starts # Data start locations (0-indexed)
#> [1] 0 1 3
letter_groups$counts # Group counts
#> [1] 7 6 7
letter_groups$order # Order permutation (0-indexed)
#> [1] 0 2 4 11 12 13 19 1 5 6 9 10 18 3 7 8 14 15 16 17
letter_groups$group_names # Names of unique groups
#> [1] "A" "C" "B"In real analysis, you normally wouldn’t need all of this metadata.
For example, to return a data frame of unique groups (keys) and their
associated counts, we only need starts() and
counts().
template <RVector T>
[[cppally::register]]
r_df group_counts(T x){
groups g = make_groups(x, /*ordered=*/ false);
return make_df(
arg("key") = subset(x, g.starts()),
arg("count") = g.counts()
);
}
group_counts(x)
#> key count
#> 1 A 7
#> 2 C 6
#> 3 B 7To illustrate how fast make_groups() is, let’s compare
cppally-derived group counts of a relatively large sample to base-R
derived group counts.
# 10k groups, 100k data points
x <- sample.int(10^4, 10^5, replace = TRUE)
mark(
cppally_group_counts = group_counts(x)$count,
base_table = table(factor(x, levels = unique(x))) |> as.integer()
)
#> # A tibble: 2 × 6
#> expression min median `itr/sec` mem_alloc `gc/sec`
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl>
#> 1 cppally_group_counts 227.32µs 370.56µs 2909. 514.1KB 46.1
#> 2 base_table 6.19ms 6.29ms 145. 6.2MB 35.5Annex
Unsupported classes in R-registered templates
r_sym and r_function are both unsupported
in R-registered template functions. This only applies when used as
template arguments. They will work as normal where function arguments
are non-templated.
Good:
Bad:
All core cppally concepts
RNumber - Includes
r_int,r_int64andr_dblRIntegerType - Includes
r_lgl,r_int,r_int64RIntegerNumber - Includes
r_intandr_int64RMathType - Includes
r_lgl,r_int,r_int64andr_dblRStringType - Includes
r_strandr_str_viewRScalar - Includes
r_lgl,r_int,r_int64,r_dbl,r_str,r_str_view,r_cplx,r_raw,r_dateandr_psxctRVal - Includes anything a cppally vector (
r_vector<>) can contain: RScalar +r_sexpRVector - Includes
r_vector<T>whereTis an RValRFactor - Factors
RDataFrame - Data frames
RComposite - Includes vectors, factors and data frames
RTimeType - Includes
r_dateandr_psxctRNumericType - Numeric types, including RMathType and RTimeType
RSortableType - Includes RNumericType and RStringType (strings can also be sorted)
RAtomicVector - A vector that contains RScalar elements
CppallyType - Any R type defined by R, including RVal, RVector, RFactor, RDataFrame, RSymbol
CppType - Anything that is not an CppallyType
CastableToRScalar - Anything that can be constructed or cast into an RScalar (which also includes RScalar)
Other useful type traits
-
unwrap_t- Returns the underlying unwrapped type -
as_r_scalar_t- Returns the equivalent RScalar type -
as_r_composite_t- Returns the equivalent RComposite type -
common_r_t- Returns the common cppally type between 2 types
Accessing the underlying types and values
While it is generally not recommended to access the underlying
values, you can do so with unwrap() which returns the
underlying C/C++ value. For example, unwrap(r_int(5)) will
return an int of value 5.
To access the underlying type, use unwrap_t<>
which always aligns with unwrap()
int one = unwrap(r_int(1));
double two = unwrap(r_dbl(2));
// Confirming that unwrap_t produces the required types as well
static_assert(std::is_same_v<int, unwrap_t<r_int>>);
static_assert(std::is_same_v<double, unwrap_t<r_dbl>>);
// unwrap_t<r_str> = SEXP
// cppally recommends against using BOTH the cppally API and the R C API together.
unwrap_t<r_str> r_c_api_blank_str = unwrap(r_str(""));