galaaz 2.1.7 → 2.1.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +9 -0
- data/blogs/galaaz_ggplot/galaaz_ggplot.Rmd +63 -58
- data/blogs/galaaz_ggplot/galaaz_ggplot.log +59 -68
- data/blogs/galaaz_ggplot/galaaz_ggplot.md +91 -84
- data/blogs/galaaz_ggplot/galaaz_ggplot.tex +125 -94
- data/blogs/galaaz_ggplot/galaaz_ggplot_files/figure-html/midwest_rb.png +0 -0
- data/blogs/galaaz_ggplot/galaaz_ggplot_files/figure-html/scatter_plot_rb.png +0 -0
- data/blogs/galaaz_ggplot/galaaz_ggplot_files/figure-markdown_github/midwest_rb.png +0 -0
- data/blogs/galaaz_ggplot/galaaz_ggplot_files/figure-markdown_github/scatter_plot_rb.png +0 -0
- data/blogs/gknit/gknit.Rmd +33 -28
- data/blogs/gknit/gknit.md +47 -42
- data/blogs/gknit/gknit.tex +1368 -0
- data/blogs/gknit/gknit_files/figure-html/bubble-1.png +0 -0
- data/blogs/gknit/gknit_files/figure-html/diverging_bar.png +0 -0
- data/blogs/gknit/gknit_files/figure-latex/bubble-1.png +0 -0
- data/blogs/gknit/gknit_files/gknit_files/figure-latex/bubble-1.png +0 -0
- data/blogs/manual/manual.Rmd +129 -60
- data/blogs/manual/manual.log +289 -545
- data/blogs/manual/manual.md +551 -467
- data/blogs/manual/manual.tex +1059 -485
- data/blogs/manual/manual_files/figure-html/bubble-1.png +0 -0
- data/blogs/manual/manual_files/figure-latex/bubble-1.png +0 -0
- data/blogs/manual/manual_files/figure-markdown_github/bubble-1.png +0 -0
- data/blogs/manual/manual_files/figure-markdown_github/diverging_bar.png +0 -0
- data/blogs/manual/manual_files/manual_files/figure-latex/bubble-1.png +0 -0
- data/blogs/nse_dplyr/nse_dplyr.Rmd +28 -7
- data/blogs/nse_dplyr/nse_dplyr.log +49 -153
- data/blogs/nse_dplyr/nse_dplyr.md +676 -705
- data/blogs/nse_dplyr/nse_dplyr.tex +1589 -0
- data/blogs/oh_my/oh_my.Rmd +193 -55
- data/blogs/oh_my/oh_my.log +265 -95
- data/blogs/oh_my/oh_my.md +236 -95
- data/blogs/oh_my/oh_my.tex +1976 -68
- data/blogs/ruby_plot/ruby_plot.Rmd +42 -34
- data/blogs/ruby_plot/ruby_plot.log +101 -99
- data/blogs/ruby_plot/ruby_plot.md +52 -46
- data/blogs/ruby_plot/ruby_plot.tex +134 -102
- data/blogs/ruby_plot/ruby_plot_files/figure-html/dose_len.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/facet_by_delivery.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/facet_by_dose.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_by_delivery_color.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_by_delivery_color2.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_with_decorations.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_with_points.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/final_box_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/final_violin_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/violin_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/dose_len.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/facet_by_delivery.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/facet_by_dose.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_by_delivery_color.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_by_delivery_color2.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_with_decorations.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_with_points.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/final_box_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/final_violin_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/violin_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/dose_len.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facet_by_delivery.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facet_by_dose.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_by_delivery_color.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_by_delivery_color2.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_with_decorations.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_with_points.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/final_box_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/final_violin_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/violin_with_jitter.png +0 -0
- data/lib/galaaz/cli.rb +51 -10
- data/script/omarchy/README.md +1 -1
- data/script/omarchy/galaaz-guide.sh +1 -1
- data/sty/galaaz.sty +22 -0
- data/version.rb +1 -1
- metadata +19 -1
|
@@ -1,140 +1,134 @@
|
|
|
1
|
-
---
|
|
2
|
-
title: "Non Standard Evaluation in dplyr with Galaaz"
|
|
3
|
-
author:
|
|
4
|
-
- "Rodrigo Botafogo"
|
|
5
|
-
- "Daniel Mossé - University of Pittsburgh"
|
|
6
|
-
tags: [Tech, Data Science, Ruby, R, JRuby, "GNU R", Galaaz, dplyr]
|
|
7
|
-
date: "10/05/2019 (narrative updated for Galaaz 2.0, 2026)"
|
|
8
|
-
output:
|
|
9
|
-
html_document:
|
|
10
|
-
self_contained: true
|
|
11
|
-
keep_md: true
|
|
12
|
-
pdf_document:
|
|
13
|
-
includes:
|
|
14
|
-
in_header: ["../../sty/galaaz.sty"]
|
|
15
|
-
number_sections: yes
|
|
16
|
-
toc: true
|
|
17
|
-
toc_depth: 2
|
|
18
|
-
md_document:
|
|
19
|
-
variant: markdown_github
|
|
20
|
-
fontsize: 11pt
|
|
21
|
-
---
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
1
|
# Introduction
|
|
26
2
|
|
|
27
|
-
According to Steven Sagaert’s answer on Quora about “Is programming
|
|
3
|
+
According to Steven Sagaert’s answer on Quora about “Is programming
|
|
4
|
+
language R overrated?”:
|
|
28
5
|
|
|
29
|
-
> R is a sophisticated language with an unusual (i.e.
|
|
30
|
-
> an impure functional programming language with
|
|
31
|
-
> different OO systems.
|
|
6
|
+
> R is a sophisticated language with an unusual (i.e. non-mainstream)
|
|
7
|
+
> set of features. It‘s an impure functional programming language with
|
|
8
|
+
> sophisticated metaprogramming and 3 different OO systems.
|
|
32
9
|
|
|
33
|
-
> Just like common lisp you can completely customise how things work via
|
|
34
|
-
> The biggest example is the tidyverse: by creating
|
|
35
|
-
> was able to create a custom
|
|
10
|
+
> Just like common lisp you can completely customise how things work via
|
|
11
|
+
> metaprogramming. The biggest example is the tidyverse: by creating
|
|
12
|
+
> it’s own evaluation system (tidyeval) was able to create a custom
|
|
13
|
+
> syntax for dplyr.
|
|
36
14
|
|
|
37
|
-
> Mastering R (the language) and its ecosystem is not a matter of weeks
|
|
38
|
-
> takes years. The rabbit hole goes pretty deep…
|
|
15
|
+
> Mastering R (the language) and its ecosystem is not a matter of weeks
|
|
16
|
+
> or months but takes years. The rabbit hole goes pretty deep…
|
|
39
17
|
|
|
40
|
-
Although a highly configurable language can give programmers a great
|
|
41
|
-
it can also take years to master—as noted above.
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
18
|
+
Although a highly configurable language can give programmers a great
|
|
19
|
+
deal of power, it can also take years to master—as noted above.
|
|
20
|
+
Programming with *dplyr*, for instance, means learning evaluation rules
|
|
21
|
+
that are not always approachable for **statisticians and analysts who
|
|
22
|
+
are not full-time software engineers**. That is not a criticism: R was
|
|
23
|
+
**built** for **statisticians** who need trustworthy results on a
|
|
24
|
+
deadline, not necessarily for building large applications.
|
|
46
25
|
|
|
47
|
-
**Unfortunately**, when such a user moves on to more **sophisticated**
|
|
48
|
-
the learning curve can become a real hurdle.
|
|
26
|
+
**Unfortunately**, when such a user moves on to more **sophisticated**
|
|
27
|
+
programming patterns, the learning curve can become a real hurdle.
|
|
49
28
|
|
|
50
|
-
In this post we will see how to program with
|
|
51
|
-
the learning curve of mastering
|
|
29
|
+
In this post we will see how to program with *dplyr* in Galaaz and how
|
|
30
|
+
Ruby can simplify the learning curve of mastering *dplyr* coding.
|
|
52
31
|
|
|
53
32
|
# But first, what is Galaaz??
|
|
54
33
|
|
|
55
|
-
Galaaz is a system for tightly coupling Ruby and R.
|
|
56
|
-
a large community, a very large set of libraries and
|
|
57
|
-
easy to learn. However,
|
|
58
|
-
|
|
59
|
-
On the other hand, R is considered one of the most powerful
|
|
60
|
-
above problems.
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
34
|
+
Galaaz is a system for tightly coupling Ruby and R. Ruby is a powerful
|
|
35
|
+
language, with a large community, a very large set of libraries and
|
|
36
|
+
great for web development. It is also easy to learn. However, it lacks
|
|
37
|
+
libraries for data science, statistics, scientific plotting and machine
|
|
38
|
+
learning. On the other hand, R is considered one of the most powerful
|
|
39
|
+
languages for solving all of the above problems. **Python** is a strong
|
|
40
|
+
competitor, with NumPy, pandas, SciPy, scikit-learn, and **many
|
|
41
|
+
thousands** of other packages on PyPI. We will not dwell on R **versus**
|
|
42
|
+
Python here: both are excellent languages with different strengths. Our
|
|
43
|
+
interest is to bring to yet another excellent language, Ruby, the data
|
|
44
|
+
science libraries that it lacks.
|
|
45
|
+
|
|
46
|
+
With Galaaz we do not intend to re-implement any of the scientific
|
|
47
|
+
libraries in R. However, we allow for very tight coupling between the
|
|
48
|
+
two languages to the point that the Ruby developer does not need to know
|
|
49
|
+
that there is an R engine running. Also, from the point of view of the R
|
|
50
|
+
user/developer, Galaaz looks a lot like R, with just minor syntactic
|
|
51
|
+
difference, so there is almost no learning curve for the R developer.
|
|
52
|
+
And as we will see in this post that programming with *dplyr* is easier
|
|
53
|
+
in Galaaz than in R.
|
|
54
|
+
|
|
55
|
+
R users are probably quite knowledgeable about *dplyr*. For the Ruby
|
|
56
|
+
developer, *dplyr* and the *tidyverse* libraries are a set of libraries
|
|
57
|
+
for data manipulation in R, developed by Hadley Wickham, Chief Scientist
|
|
58
|
+
at Posit (formerly RStudio) and a prolific R coder and writer.
|
|
59
|
+
|
|
60
|
+
For the coupling of Ruby and R, **Galaaz 2.0** uses
|
|
61
|
+
**[JRuby](https://www.jruby.org/)** (Ruby on the JVM) together with
|
|
62
|
+
**GNU R**. A **bridge** sends expressions and data between Ruby and an R
|
|
63
|
+
process so that Ruby can call **dplyr** and the rest of the tidyverse as
|
|
64
|
+
if they were part of the same workflow. An **earlier** Galaaz line of
|
|
65
|
+
work used Oracle’s **GraalVM** with **TruffleRuby** and **FastR** in a
|
|
66
|
+
single runtime; that approach is **no longer** the supported stack—see
|
|
67
|
+
the project **manual** for setup, **`bin/galaaz-jruby`**, and **gKnit**.
|
|
84
68
|
|
|
85
69
|
# Tidyverse and dplyr
|
|
86
70
|
|
|
87
|
-
In [What is the
|
|
88
|
-
tidyverse
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
>
|
|
92
|
-
>
|
|
93
|
-
>
|
|
94
|
-
>
|
|
95
|
-
>
|
|
96
|
-
> that
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
>
|
|
105
|
-
>
|
|
106
|
-
|
|
107
|
-
>
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
multiple
|
|
71
|
+
In [What is the
|
|
72
|
+
tidyverse?](https://rviews.rstudio.com/2017/06/08/what-is-the-tidyverse/)
|
|
73
|
+
the tidyverse is explained as follows:
|
|
74
|
+
|
|
75
|
+
> The tidyverse is a coherent system of packages for data manipulation,
|
|
76
|
+
> exploration and visualization that share a common design philosophy.
|
|
77
|
+
> These were mostly developed by Hadley Wickham himself, but they are
|
|
78
|
+
> now being expanded by several contributors. Tidyverse packages are
|
|
79
|
+
> intended to make statisticians and data scientists more productive by
|
|
80
|
+
> guiding them through workflows that facilitate communication, and
|
|
81
|
+
> result in reproducible work products. Fundamentally, the tidyverse is
|
|
82
|
+
> about the connections between the tools that make the workflow
|
|
83
|
+
> possible.
|
|
84
|
+
|
|
85
|
+
*dplyr* is one of the many packages that are part of the tidyverse. It
|
|
86
|
+
is:
|
|
87
|
+
|
|
88
|
+
> a grammar of data manipulation, providing a consistent set of verbs
|
|
89
|
+
> that help you solve the most common data manipulation challenges:
|
|
90
|
+
|
|
91
|
+
> 1. mutate() adds new variables that are functions of existing
|
|
92
|
+
> variables
|
|
93
|
+
> 2. select() picks variables based on their names.
|
|
94
|
+
> 3. filter() picks cases based on their values.
|
|
95
|
+
> 4. summarise() reduces multiple values down to a single summary.
|
|
96
|
+
> 5. arrange() changes the ordering of the rows.
|
|
97
|
+
|
|
98
|
+
Very often R is used interactively and users use *dplyr* to manipulate a
|
|
99
|
+
single dataset without programming. When users want to replicate their
|
|
100
|
+
work for multiple datasets, programming becomes necessary.
|
|
112
101
|
|
|
113
102
|
# Programming with dplyr
|
|
114
103
|
|
|
115
|
-
In the vignette [
|
|
116
|
-
Hadley
|
|
104
|
+
In the vignette [“Programming with
|
|
105
|
+
dplyr”](https://dplyr.tidyverse.org/articles/programming.html), Hadley
|
|
106
|
+
Wickham states:
|
|
117
107
|
|
|
118
|
-
> Most dplyr functions use non-standard evaluation (NSE). This is a
|
|
119
|
-
> means they don’t follow the usual R rules of
|
|
120
|
-
> expression that you typed and
|
|
121
|
-
> benefits for dplyr
|
|
108
|
+
> Most dplyr functions use non-standard evaluation (NSE). This is a
|
|
109
|
+
> catch-all term that means they don’t follow the usual R rules of
|
|
110
|
+
> evaluation. Instead, they capture the expression that you typed and
|
|
111
|
+
> evaluate it in a custom way. This has two main benefits for dplyr
|
|
112
|
+
> code:
|
|
122
113
|
|
|
123
|
-
> Operations on data frames can be expressed succinctly because you
|
|
124
|
-
> the name of the data frame. For example, you can
|
|
125
|
-
>
|
|
114
|
+
> Operations on data frames can be expressed succinctly because you
|
|
115
|
+
> don’t need to repeat the name of the data frame. For example, you can
|
|
116
|
+
> write filter(df, x == 1, y == 2, z == 3) instead of df\[df$x == 1 &
|
|
117
|
+
> df$y ==2 & df$z == 3, \].
|
|
126
118
|
|
|
127
|
-
> dplyr can choose to compute results in a different way to base R. This
|
|
128
|
-
> database backends because dplyr itself doesn’t do any
|
|
129
|
-
> that tells the database what to
|
|
119
|
+
> dplyr can choose to compute results in a different way to base R. This
|
|
120
|
+
> is important for database backends because dplyr itself doesn’t do any
|
|
121
|
+
> work, but instead generates the SQL that tells the database what to
|
|
122
|
+
> do.
|
|
130
123
|
|
|
131
124
|
But then he goes on:
|
|
132
125
|
|
|
133
|
-
> Unfortunately these benefits do not come for free. There are two main
|
|
134
|
-
|
|
135
|
-
> Most dplyr arguments are not referentially transparent. That means you can’t replace a value
|
|
136
|
-
> with a seemingly equivalent object that you’ve defined elsewhere. In other words, this code:
|
|
126
|
+
> Unfortunately these benefits do not come for free. There are two main
|
|
127
|
+
> drawbacks:
|
|
137
128
|
|
|
129
|
+
> Most dplyr arguments are not referentially transparent. That means you
|
|
130
|
+
> can’t replace a value with a seemingly equivalent object that you’ve
|
|
131
|
+
> defined elsewhere. In other words, this code:
|
|
138
132
|
|
|
139
133
|
``` r
|
|
140
134
|
df <- data.frame(x = 1:3, y = 3:1)
|
|
@@ -144,8 +138,8 @@ print(filter(df, x == 1))
|
|
|
144
138
|
#> <int> <int>
|
|
145
139
|
#> 1 1 3
|
|
146
140
|
```
|
|
147
|
-
> Is not equivalent to this code:
|
|
148
141
|
|
|
142
|
+
> Is not equivalent to this code:
|
|
149
143
|
|
|
150
144
|
``` r
|
|
151
145
|
my_var <- x
|
|
@@ -153,109 +147,110 @@ my_var <- x
|
|
|
153
147
|
filter(df, my_var == 1)
|
|
154
148
|
#> Error: object 'my_var' not found
|
|
155
149
|
```
|
|
156
|
-
> This makes it hard to create functions with arguments that change how dplyr verbs are computed.
|
|
157
150
|
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
151
|
+
> This makes it hard to create functions with arguments that change how
|
|
152
|
+
> dplyr verbs are computed.
|
|
153
|
+
|
|
154
|
+
As a result of this, programming with *dplyr* requires learning a set of
|
|
155
|
+
new ideas and concepts. In this vignette Hadley goes on showing how to
|
|
156
|
+
program ever more difficult problems with *dplyr*, showing the problems
|
|
157
|
+
it faces and the new concepts needed to solve them.
|
|
161
158
|
|
|
162
|
-
In this blog, we will look at all the problems presented by Harley on
|
|
163
|
-
those same problems can be solved using Galaaz
|
|
159
|
+
In this blog, we will look at all the problems presented by Harley on
|
|
160
|
+
the vignette and show how those same problems can be solved using Galaaz
|
|
161
|
+
and the Ruby language.
|
|
164
162
|
|
|
165
|
-
This blog is organized as follows: first we show how to write
|
|
166
|
-
|
|
167
|
-
|
|
163
|
+
This blog is organized as follows: first we show how to write
|
|
164
|
+
expressions using Galaaz.
|
|
165
|
+
Expressions are a fundamental concept in *dplyr* and are not part of
|
|
166
|
+
basic Ruby. We extend the Ruby language create a manipulate expressions
|
|
167
|
+
that will be used by *dplyr* functions.
|
|
168
168
|
|
|
169
|
-
Then we show very succintly how Ruby and R can be integrated and how R
|
|
170
|
-
transparently called from Ruby. Galaaz [user
|
|
171
|
-
(still in development)
|
|
169
|
+
Then we show very succintly how Ruby and R can be integrated and how R
|
|
170
|
+
functions are transparently called from Ruby. Galaaz [user
|
|
171
|
+
manual](https://github.com/rbotafogo/galaaz/wiki) (still in development)
|
|
172
|
+
goes in much deeper detail about this integration.
|
|
172
173
|
|
|
173
|
-
Next in section
|
|
174
|
-
|
|
175
|
-
|
|
174
|
+
Next in section “Data manipulation wiht *dplyr*” we go through all the
|
|
175
|
+
problems on the *dplyr* vignette and look at how they are solved in
|
|
176
|
+
Galaaz. We then discuss why programming with Galaaz and *dplyr* is
|
|
177
|
+
easier than programming with *dplyr* in plain R.
|
|
176
178
|
|
|
177
|
-
The following section looks at another more advanced problem and shows
|
|
178
|
-
handle it without any difficulty.
|
|
179
|
+
The following section looks at another more advanced problem and shows
|
|
180
|
+
that Galaaz can still handle it without any difficulty. We then provide
|
|
181
|
+
further reading and concluding remarks.
|
|
179
182
|
|
|
180
183
|
# Writing Expressions in Galaaz
|
|
181
184
|
|
|
182
|
-
Galaaz extends Ruby to work with expressions, similar to R
|
|
183
|
-
(base R) or
|
|
184
|
-
|
|
185
|
-
|
|
185
|
+
Galaaz extends Ruby to work with expressions, similar to R’s expressions
|
|
186
|
+
build with ‘quote’ (base R) or ‘quo’ (tidyverse). Expressions in this
|
|
187
|
+
context are like mathematical expressions or formulae. For instance, in
|
|
188
|
+
mathematics, the expression *y* = *s**i**n*(*x*) describes a function
|
|
189
|
+
but cannot be computed unless the value of *x* is bound to some value.
|
|
186
190
|
|
|
187
|
-
Expressions are fundamental in
|
|
188
|
-
for instance, as we will see shortly, if a data
|
|
189
|
-
|
|
190
|
-
|
|
191
|
+
Expressions are fundamental in *dplyr* programming as they are the input
|
|
192
|
+
to *dplyr* functions, for instance, as we will see shortly, if a data
|
|
193
|
+
frame has a column named ‘x’ and we want to add another column, y, to
|
|
194
|
+
this dataframe that has the values of ‘x’ times 2, then we would call a
|
|
195
|
+
*dplyr* function with the expression ‘y = x \* 2’.
|
|
191
196
|
|
|
192
197
|
## A note on notation
|
|
193
198
|
|
|
194
|
-
This blog was written in Rmarkdown and automatically converted to HTML
|
|
195
|
-
where you are reading this blog) with gKnit (a tool
|
|
196
|
-
|
|
197
|
-
blocks
|
|
198
|
-
|
|
199
|
-
|
|
199
|
+
This blog was written in Rmarkdown and automatically converted to HTML
|
|
200
|
+
or PDF (depending on where you are reading this blog) with gKnit (a tool
|
|
201
|
+
provided by Galaaz). In Rmarkdown, it is possible to write text and code
|
|
202
|
+
blocks that are executed to generate the final report. Code blocks
|
|
203
|
+
appear inside a ‘box’ and the result of their execution appear either in
|
|
204
|
+
another type of ‘box’ with a different background (HTML) or as normal
|
|
205
|
+
text (PDF). Every output line from the code execution is preceded by
|
|
206
|
+
‘\##’.
|
|
200
207
|
|
|
201
208
|
## Expressions from operators
|
|
202
209
|
|
|
203
|
-
The code below creates an expression summing two symbols. Note that :a
|
|
204
|
-
are not bound to any values at the time of
|
|
205
|
-
|
|
210
|
+
The code below creates an expression summing two symbols. Note that :a
|
|
211
|
+
and :b are Ruby symbols and are not bound to any values at the time of
|
|
212
|
+
expression definition:
|
|
206
213
|
|
|
207
214
|
``` ruby
|
|
208
|
-
|
|
209
|
-
|
|
215
|
+
begin
|
|
216
|
+
exp1 = :a + :b
|
|
217
|
+
puts exp1
|
|
218
|
+
rescue => e
|
|
219
|
+
# Bare Symbol#+ is not expression sugar in Galaaz 2.0; short error for PDF.
|
|
220
|
+
puts "#{e.class}: #{e.message}"
|
|
221
|
+
end
|
|
210
222
|
```
|
|
211
223
|
|
|
212
|
-
|
|
213
|
-
## undefined method '+' for an instance of Symbol
|
|
214
|
-
```
|
|
224
|
+
## NoMethodError: undefined method '+' for an instance of Symbol
|
|
215
225
|
|
|
216
|
-
```
|
|
217
|
-
## /home/rbotafogo/desenv_linux/galaaz/lib/util/exec_ruby.rb:170:in 'exec_ruby'
|
|
218
|
-
## org/jruby/RubyKernel.java:1268:in 'eval'
|
|
219
|
-
## /home/rbotafogo/desenv_linux/galaaz/lib/util/exec_ruby.rb:169:in 'exec_ruby'
|
|
220
|
-
## /home/rbotafogo/desenv_linux/galaaz/lib/gknit/knitr_engine.rb:777:in 'block in initialize'
|
|
221
|
-
## org/jruby/RubyBasicObject.java:2695:in 'instance_eval'
|
|
222
|
-
## org/jruby/RubyBasicObject.java:2723:in 'instance_eval'
|
|
223
|
-
## /home/rbotafogo/desenv_linux/galaaz/lib/gknit/knitr_engine.rb:748:in 'block in initialize'
|
|
224
|
-
## /home/rbotafogo/desenv_linux/galaaz/lib/R_interface/new_bridge_adapter.rb:358:in 'block in register_callback_proc_stub'
|
|
225
|
-
## /home/rbotafogo/desenv_linux/galaaz/lib/new_bridge/session_client.rb:413:in 'block in handle_call'
|
|
226
|
-
```
|
|
227
226
|
In Galaaz, we can build any complex mathematical expression such as:
|
|
228
227
|
|
|
229
|
-
|
|
230
228
|
``` ruby
|
|
231
229
|
exp2 = (R[:a] + R[:b]) * 2.0 + R[:c] ** 2 / R[:z]
|
|
232
230
|
puts exp2
|
|
233
231
|
```
|
|
234
232
|
|
|
235
|
-
|
|
236
|
-
## a + b * 2.0 + c ^ 2L / z
|
|
237
|
-
```
|
|
238
|
-
Expressions are printed with the same format as the equivalent R expressions. The 'L' after
|
|
239
|
-
2 indicates that 2 is an integer.
|
|
233
|
+
## a + b * 2.0 + c ^ 2L / z
|
|
240
234
|
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
should write '2L'. Galaaz follows Ruby notation and '2' is an integer, while '2.0' is a
|
|
244
|
-
float.
|
|
235
|
+
Expressions are printed with the same format as the equivalent R
|
|
236
|
+
expressions. The ‘L’ after 2 indicates that 2 is an integer.
|
|
245
237
|
|
|
246
|
-
|
|
238
|
+
The R developer should note that in R, if she writes the number ‘2’, the
|
|
239
|
+
R interpreter will convert it to float. In order to get an interger she
|
|
240
|
+
should write ‘2L’. Galaaz follows Ruby notation and ‘2’ is an integer,
|
|
241
|
+
while ‘2.0’ is a float.
|
|
247
242
|
|
|
243
|
+
It is also possible to use inequality operators in building expressions:
|
|
248
244
|
|
|
249
245
|
``` ruby
|
|
250
246
|
exp3 = (R[:a] + R[:b]) >= R[:z]
|
|
251
247
|
puts exp3
|
|
252
248
|
```
|
|
253
249
|
|
|
254
|
-
|
|
255
|
-
## a + b >= z
|
|
256
|
-
```
|
|
257
|
-
Expressions' definition can also make use of normal Ruby variables without any problem:
|
|
250
|
+
## a + b >= z
|
|
258
251
|
|
|
252
|
+
Expressions’ definition can also make use of normal Ruby variables
|
|
253
|
+
without any problem:
|
|
259
254
|
|
|
260
255
|
``` ruby
|
|
261
256
|
x = 20
|
|
@@ -264,133 +259,116 @@ exp_var = (R[:a] + R[:b]) * x <= R[:z] - y
|
|
|
264
259
|
puts exp_var
|
|
265
260
|
```
|
|
266
261
|
|
|
267
|
-
|
|
268
|
-
## a + b * 20L <= z - 30.0
|
|
269
|
-
```
|
|
270
|
-
|
|
271
|
-
Galaaz provides both symbolic representations for operators, such as (>, <, !=) as functional
|
|
272
|
-
notation for those operators such as (.gt, .ge, etc.). So the same expression written
|
|
273
|
-
above can also be written as
|
|
262
|
+
## a + b * 20L <= z - 30.0
|
|
274
263
|
|
|
264
|
+
Galaaz provides both symbolic representations for operators, such as
|
|
265
|
+
(\>, \<, !=) as functional notation for those operators such as (.gt,
|
|
266
|
+
.ge, etc.). So the same expression written above can also be written as
|
|
275
267
|
|
|
276
268
|
``` ruby
|
|
277
269
|
exp4 = (R[:a] + R[:b]).ge R[:z]
|
|
278
270
|
puts exp4
|
|
279
271
|
```
|
|
280
272
|
|
|
281
|
-
|
|
282
|
-
## a + b >= z
|
|
283
|
-
```
|
|
273
|
+
## a + b >= z
|
|
284
274
|
|
|
285
|
-
Two types of expressions, however, can only be created with the
|
|
286
|
-
of the operators.
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
In order to write an expression involving '==' we
|
|
290
|
-
need to use the method '.eq' and for '=' we need the function '.assign':
|
|
275
|
+
Two types of expressions, however, can only be created with the
|
|
276
|
+
functional representation of the operators. Those are expressions
|
|
277
|
+
involving ‘==’, and ‘=’. This is the case since those symbols have
|
|
278
|
+
special meaning in Ruby and should not be redefined.
|
|
291
279
|
|
|
280
|
+
In order to write an expression involving ‘==’ we need to use the method
|
|
281
|
+
‘.eq’ and for ‘=’ we need the function ‘.assign’:
|
|
292
282
|
|
|
293
283
|
``` ruby
|
|
294
284
|
exp5 = (R[:a] + R[:b]).eq R[:z]
|
|
295
285
|
puts exp5
|
|
296
286
|
```
|
|
297
287
|
|
|
298
|
-
|
|
299
|
-
## a + b == z
|
|
300
|
-
```
|
|
301
|
-
|
|
288
|
+
## a + b == z
|
|
302
289
|
|
|
303
290
|
``` ruby
|
|
304
291
|
exp6 = R[:y].assign R[:a] + R[:b]
|
|
305
292
|
puts exp6
|
|
306
293
|
```
|
|
307
294
|
|
|
308
|
-
|
|
309
|
-
## y <- a + b
|
|
310
|
-
```
|
|
311
|
-
Users should be careful when writing expressions not to inadvertently use '==' or '=' as
|
|
312
|
-
this will generate an error, that might be a bit cryptic (in future releases of Galaza, we
|
|
313
|
-
plan to improve the error message).
|
|
295
|
+
## y <- a + b
|
|
314
296
|
|
|
297
|
+
Users should be careful when writing expressions not to inadvertently
|
|
298
|
+
use ‘==’ or ‘=’ as this will generate an error, that might be a bit
|
|
299
|
+
cryptic (in future releases of Galaza, we plan to improve the error
|
|
300
|
+
message).
|
|
315
301
|
|
|
316
302
|
``` ruby
|
|
317
303
|
exp_wrong = (R[:a] + R[:b]) == R[:z]
|
|
318
304
|
puts exp_wrong
|
|
319
305
|
```
|
|
320
306
|
|
|
321
|
-
|
|
322
|
-
## false
|
|
323
|
-
```
|
|
324
|
-
The problem lies with the fact that
|
|
325
|
-
when using '==' we are comparing expression (R[:a] + R[:b]) to expression R[:z] with '=='. When this
|
|
326
|
-
comparison is executed, the system tries to evaluate :a, :b and :z, and those symbols, at
|
|
327
|
-
this time, are not bound to anything giving the "object 'a' not found" message.
|
|
307
|
+
## false
|
|
328
308
|
|
|
329
|
-
|
|
309
|
+
The problem lies with the fact that when using ‘==’ we are comparing
|
|
310
|
+
expression (R\[:a\] + R\[:b\]) to expression R\[:z\] with ‘==’. When
|
|
311
|
+
this comparison is executed, the system tries to evaluate :a, :b and :z,
|
|
312
|
+
and those symbols, at this time, are not bound to anything giving the
|
|
313
|
+
“object ‘a’ not found” message.
|
|
330
314
|
|
|
331
|
-
|
|
332
|
-
mathematics, it's quite natural to write an expressin such as $y = sin(x)$. In this case, the
|
|
333
|
-
'sin' function is part of the expression and should not be immediately executed. When we want
|
|
334
|
-
the function to be part of the expression, we call the function preceeding it
|
|
335
|
-
by the letter E, such as 'E.sin(x)'
|
|
315
|
+
## Expressions with R methods
|
|
336
316
|
|
|
317
|
+
It is often necessary to create an expression that uses a method or
|
|
318
|
+
function. For instance, in mathematics, it’s quite natural to write an
|
|
319
|
+
expressin such as *y* = *s**i**n*(*x*). In this case, the ‘sin’ function
|
|
320
|
+
is part of the expression and should not be immediately executed. When
|
|
321
|
+
we want the function to be part of the expression, we call the function
|
|
322
|
+
preceeding it by the letter E, such as ‘E.sin(x)’
|
|
337
323
|
|
|
338
324
|
``` ruby
|
|
339
325
|
exp7 = R[:y].assign E.sin(R[:x])
|
|
340
326
|
puts exp7
|
|
341
327
|
```
|
|
342
328
|
|
|
343
|
-
|
|
344
|
-
## y <- sin(x)
|
|
345
|
-
```
|
|
346
|
-
Function expressions can also be written using '.' notation:
|
|
329
|
+
## y <- sin(x)
|
|
347
330
|
|
|
331
|
+
Function expressions can also be written using ‘.’ notation:
|
|
348
332
|
|
|
349
333
|
``` ruby
|
|
350
334
|
exp8 = R[:y].assign R[:x].sin
|
|
351
335
|
puts exp8
|
|
352
336
|
```
|
|
353
337
|
|
|
354
|
-
|
|
355
|
-
## y <- sin(x)
|
|
356
|
-
```
|
|
357
|
-
When a function has multiple arguments, the first one can be used before the '.'. For instance,
|
|
358
|
-
the R concatenate function 'c', that concatenates two or more arguments can be part of
|
|
359
|
-
an expression as:
|
|
338
|
+
## y <- sin(x)
|
|
360
339
|
|
|
340
|
+
When a function has multiple arguments, the first one can be used before
|
|
341
|
+
the ‘.’. For instance, the R concatenate function ‘c’, that concatenates
|
|
342
|
+
two or more arguments can be part of an expression as:
|
|
361
343
|
|
|
362
344
|
``` ruby
|
|
363
345
|
exp9 = R[:x].c(R[:y])
|
|
364
346
|
puts exp9
|
|
365
347
|
```
|
|
366
348
|
|
|
367
|
-
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
pipe.
|
|
349
|
+
## c(x, y)
|
|
350
|
+
|
|
351
|
+
Note that this gives an OO feeling to the code, as if we were saying ‘x’
|
|
352
|
+
concatenates ‘y’. As a side note, ‘.’ notation can be used as the R pipe
|
|
353
|
+
operator ‘%\>%’, but is more general than the pipe.
|
|
373
354
|
|
|
374
355
|
## Evaluating an Expression
|
|
375
356
|
|
|
376
|
-
Although we are mainly focusing on expressions to pass them to
|
|
377
|
-
can be evaluated by calling function
|
|
357
|
+
Although we are mainly focusing on expressions to pass them to *dplyr*
|
|
358
|
+
functions, expressions can be evaluated by calling function ‘eval’ with
|
|
359
|
+
a binding.
|
|
378
360
|
|
|
379
361
|
A binding can be provided with a list or a data frame as shown below:
|
|
380
362
|
|
|
381
|
-
|
|
382
363
|
``` ruby
|
|
383
364
|
exp = (R[:a] + R[:b]) * 2.0 + R[:c] ** 2 / R[:z]
|
|
384
365
|
puts exp.eval(R.list(a: 10, b: 20, c: 30, z: 40))
|
|
385
366
|
```
|
|
386
367
|
|
|
387
|
-
|
|
388
|
-
## [1] 72.5
|
|
389
|
-
```
|
|
368
|
+
## [1] 72.5
|
|
390
369
|
|
|
391
370
|
with a data frame:
|
|
392
371
|
|
|
393
|
-
|
|
394
372
|
``` ruby
|
|
395
373
|
df = R.data__frame(
|
|
396
374
|
a: R.c(1, 2, 3),
|
|
@@ -401,90 +379,83 @@ df = R.data__frame(
|
|
|
401
379
|
puts exp.eval(df)
|
|
402
380
|
```
|
|
403
381
|
|
|
404
|
-
|
|
405
|
-
## [1] 31 62 93
|
|
406
|
-
```
|
|
382
|
+
## [1] 31 62 93
|
|
407
383
|
|
|
408
384
|
# Using Galaaz to call R functions
|
|
409
385
|
|
|
410
|
-
Galaaz tries to emulate as closely as possible the way R functions are
|
|
411
|
-
R to Galaaz should be quite easy requiring
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
|
|
416
|
-
Basically, to call an R function from Ruby with Galaaz, one only needs to preced the function
|
|
417
|
-
with 'R.'. For instance, to create a vector in R, the 'c' function is used. In Galaaz, a
|
|
418
|
-
vector can be created by using 'R.c':
|
|
386
|
+
Galaaz tries to emulate as closely as possible the way R functions are
|
|
387
|
+
called and migrating from R to Galaaz should be quite easy requiring
|
|
388
|
+
only minor syntactic changes to an R script. In this post, we do not
|
|
389
|
+
have enough space to write a complete manual on Galaaz (a short manual
|
|
390
|
+
can be found at: <https://www.rubydoc.info/gems/galaaz/0.4.9>), so we
|
|
391
|
+
will present only a few examples scripts using Galaaz.
|
|
419
392
|
|
|
393
|
+
Basically, to call an R function from Ruby with Galaaz, one only needs
|
|
394
|
+
to preced the function with ‘R.’. For instance, to create a vector in R,
|
|
395
|
+
the ‘c’ function is used. In Galaaz, a vector can be created by using
|
|
396
|
+
‘R.c’:
|
|
420
397
|
|
|
421
398
|
``` ruby
|
|
422
399
|
vec = R.c(1.0, 2, 3)
|
|
423
400
|
puts vec
|
|
424
401
|
```
|
|
425
402
|
|
|
426
|
-
|
|
427
|
-
## [1] 1 2 3
|
|
428
|
-
```
|
|
429
|
-
A list is created in R with the 'list' function, so in Galaaz we do:
|
|
403
|
+
## [1] 1 2 3
|
|
430
404
|
|
|
405
|
+
A list is created in R with the ‘list’ function, so in Galaaz we do:
|
|
431
406
|
|
|
432
407
|
``` ruby
|
|
433
408
|
list = R.list(a: 1.0, b: 2, c: 3)
|
|
434
409
|
puts list
|
|
435
410
|
```
|
|
436
411
|
|
|
437
|
-
|
|
438
|
-
##
|
|
439
|
-
##
|
|
440
|
-
##
|
|
441
|
-
##
|
|
442
|
-
##
|
|
443
|
-
##
|
|
444
|
-
##
|
|
445
|
-
## [1] 3
|
|
446
|
-
```
|
|
447
|
-
Note that we can use named arguments in our list. The same code in R would be:
|
|
412
|
+
## $a
|
|
413
|
+
## [1] 1
|
|
414
|
+
##
|
|
415
|
+
## $b
|
|
416
|
+
## [1] 2
|
|
417
|
+
##
|
|
418
|
+
## $c
|
|
419
|
+
## [1] 3
|
|
448
420
|
|
|
421
|
+
Note that we can use named arguments in our list. The same code in R
|
|
422
|
+
would be:
|
|
449
423
|
|
|
450
424
|
``` r
|
|
451
425
|
lst = list(a = 1, b = 2L, c = 3L)
|
|
452
426
|
print(lst)
|
|
453
427
|
```
|
|
454
428
|
|
|
455
|
-
|
|
456
|
-
##
|
|
457
|
-
##
|
|
458
|
-
##
|
|
459
|
-
##
|
|
460
|
-
##
|
|
461
|
-
##
|
|
462
|
-
##
|
|
463
|
-
## [1] 3
|
|
464
|
-
```
|
|
465
|
-
Now, let's say that 'x' is an angle of 45$^\circ$ and we acttually want to create
|
|
466
|
-
the expression $y = sin(45^\circ)$, which is $y = 0.850...$. In this case,
|
|
467
|
-
we will use 'R.sin':
|
|
429
|
+
## $a
|
|
430
|
+
## [1] 1
|
|
431
|
+
##
|
|
432
|
+
## $b
|
|
433
|
+
## [1] 2
|
|
434
|
+
##
|
|
435
|
+
## $c
|
|
436
|
+
## [1] 3
|
|
468
437
|
|
|
438
|
+
Now, let’s say that ‘x’ is an angle of 45<sup>∘</sup> and we acttually
|
|
439
|
+
want to create the expression *y* = *s**i**n*(45<sup>∘</sup>), which is
|
|
440
|
+
*y* = 0.850.... In this case, we will use ‘R.sin’:
|
|
469
441
|
|
|
470
442
|
``` ruby
|
|
471
443
|
exp10 = R[:y].assign R.sin(45)
|
|
472
444
|
puts exp10
|
|
473
445
|
```
|
|
474
446
|
|
|
475
|
-
|
|
476
|
-
## y <- 0.850903524534118
|
|
477
|
-
```
|
|
478
|
-
|
|
479
|
-
# Data manipulation wiht _dplyr_
|
|
447
|
+
## y <- 0.850903524534118
|
|
480
448
|
|
|
481
|
-
|
|
482
|
-
data in Ruby with it. This section will follow [_dplyr_'s vignette](https://dplyr.tidyverse.org/articles/dplyr.html) that explores the nycflights13 data set. This dataset contains all 336776
|
|
483
|
-
flights that departed from New York City in 2013. The data comes from the US Bureau of
|
|
484
|
-
Transportation Statistics.
|
|
449
|
+
# Data manipulation wiht *dplyr*
|
|
485
450
|
|
|
486
|
-
|
|
451
|
+
In this section we will give a brief tour *dplyr*’s usage in Galaaz and
|
|
452
|
+
how to manipulate data in Ruby with it. This section will follow
|
|
453
|
+
[*dplyr*’s vignette](https://dplyr.tidyverse.org/articles/dplyr.html)
|
|
454
|
+
that explores the nycflights13 data set. This dataset contains all
|
|
455
|
+
336776 flights that departed from New York City in 2013. The data comes
|
|
456
|
+
from the US Bureau of Transportation Statistics.
|
|
487
457
|
|
|
458
|
+
Let’s start by taking a look at this dataset:
|
|
488
459
|
|
|
489
460
|
``` ruby
|
|
490
461
|
R.library('nycflights13')
|
|
@@ -494,140 +465,137 @@ puts ~R[:flights].dim
|
|
|
494
465
|
~R[:flights].str
|
|
495
466
|
```
|
|
496
467
|
|
|
497
|
-
|
|
498
|
-
##
|
|
499
|
-
## <environment: 0x5f0b6a5b7790>
|
|
500
|
-
```
|
|
501
|
-
|
|
502
|
-
Now, let's use a first verb of _dplyr_: 'filter'. This verb, obviously, will filter the data
|
|
503
|
-
by the given expression. In the next block, we filter by columns 'month' and 'day'. The
|
|
504
|
-
first argument to the filter function is symbol ':flights'. A Ruby symbol, when given to
|
|
505
|
-
an R function will convert to the R variable of the same name, in this case 'flights', that
|
|
506
|
-
holds the nycflights13 data frame.
|
|
468
|
+
## ~(dim(flights))
|
|
469
|
+
## <environment: 0x57d8ad019298>
|
|
507
470
|
|
|
508
|
-
|
|
509
|
-
filter
|
|
471
|
+
Now, let’s use a first verb of *dplyr*: ‘filter’. This verb, obviously,
|
|
472
|
+
will filter the data by the given expression. In the next block, we
|
|
473
|
+
filter by columns ‘month’ and ‘day’. The first argument to the filter
|
|
474
|
+
function is symbol ‘:flights’. A Ruby symbol, when given to an R
|
|
475
|
+
function will convert to the R variable of the same name, in this case
|
|
476
|
+
‘flights’, that holds the nycflights13 data frame.
|
|
510
477
|
|
|
478
|
+
The second and third arguments are expressions that will be used by the
|
|
479
|
+
filter function to filter by columns, looking for entries in which the
|
|
480
|
+
month and day are equal to 1.
|
|
511
481
|
|
|
512
482
|
``` ruby
|
|
513
483
|
puts R.filter(:flights, (R[:month].eq 1), (R[:day].eq 1))
|
|
514
484
|
```
|
|
515
485
|
|
|
516
|
-
|
|
517
|
-
##
|
|
518
|
-
##
|
|
519
|
-
##
|
|
520
|
-
##
|
|
521
|
-
##
|
|
522
|
-
##
|
|
523
|
-
##
|
|
524
|
-
##
|
|
525
|
-
##
|
|
526
|
-
##
|
|
527
|
-
##
|
|
528
|
-
##
|
|
529
|
-
##
|
|
530
|
-
## # ℹ
|
|
531
|
-
## #
|
|
532
|
-
## #
|
|
533
|
-
## #
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
537
|
-
|
|
538
|
-
|
|
539
|
-
|
|
540
|
-
|
|
541
|
-
|
|
542
|
-
this blog.
|
|
486
|
+
## # A tibble: 842 × 19
|
|
487
|
+
## year month day dep_time sched_dep_time dep_delay arr_time
|
|
488
|
+
## <int> <int> <int> <int> <int> <dbl> <int>
|
|
489
|
+
## 1 2013 1 1 517 515 2 830
|
|
490
|
+
## 2 2013 1 1 533 529 4 850
|
|
491
|
+
## 3 2013 1 1 542 540 2 923
|
|
492
|
+
## 4 2013 1 1 544 545 -1 1004
|
|
493
|
+
## 5 2013 1 1 554 600 -6 812
|
|
494
|
+
## 6 2013 1 1 554 558 -4 740
|
|
495
|
+
## 7 2013 1 1 555 600 -5 913
|
|
496
|
+
## 8 2013 1 1 557 600 -3 709
|
|
497
|
+
## 9 2013 1 1 557 600 -3 838
|
|
498
|
+
## 10 2013 1 1 558 600 -2 753
|
|
499
|
+
## # ℹ 832 more rows
|
|
500
|
+
## # ℹ 12 more variables: sched_arr_time <int>, arr_delay <dbl>,
|
|
501
|
+
## # carrier <chr>, flight <int>, tailnum <chr>, origin <chr>,
|
|
502
|
+
## # dest <chr>, air_time <dbl>, distance <dbl>, hour <dbl>,
|
|
503
|
+
## # minute <dbl>, time_hour <dttm>
|
|
504
|
+
|
|
505
|
+
## Programming with *dplyr*: problems and how to solve them in Galaaz
|
|
506
|
+
|
|
507
|
+
In this section we look at the list of problems that Hadley describes in
|
|
508
|
+
the “Programming with dplyr” vignette and show how those problems are
|
|
509
|
+
solved and coded with Galaaz. Readers interested in how those problems
|
|
510
|
+
are treated in *dplyr* should read the vignette and use it as a
|
|
511
|
+
comparison with this blog.
|
|
543
512
|
|
|
544
513
|
## Filtering using expressions
|
|
545
514
|
|
|
546
|
-
Now that we know how to write expressions and call R functions, let
|
|
547
|
-
Galaaz.
|
|
548
|
-
|
|
549
|
-
|
|
550
|
-
|
|
551
|
-
|
|
515
|
+
Now that we know how to write expressions and call R functions, let’s do
|
|
516
|
+
some data manipulation in Galaaz. Let’s first start by creating a data
|
|
517
|
+
frame. In R, the ‘data.frame’ function creates a data frame. In Ruby,
|
|
518
|
+
writing ‘data.frame’ will not parse as a single object. To call R
|
|
519
|
+
functions that have a ‘.’ in them, we need to substitute the ‘.’ with
|
|
520
|
+
’\_\_‘. So, method ’data.frame’ in R, is called in Galaaz as
|
|
521
|
+
‘R.data\_\_frame’:
|
|
552
522
|
|
|
553
523
|
``` ruby
|
|
554
524
|
df = R.data__frame(x: (1..3), y: (3..1))
|
|
555
525
|
puts df
|
|
556
526
|
```
|
|
557
527
|
|
|
558
|
-
|
|
559
|
-
##
|
|
560
|
-
##
|
|
561
|
-
##
|
|
562
|
-
## 3 3 1
|
|
563
|
-
```
|
|
564
|
-
|
|
565
|
-
_dplyr_ provides the 'filter' function, that filters data in a data brame. The 'filter'
|
|
566
|
-
function can be called on this data frame either by using 'R.filter(df, ...)' or
|
|
567
|
-
by using dot notation.
|
|
528
|
+
## x y
|
|
529
|
+
## 1 1 3
|
|
530
|
+
## 2 2 2
|
|
531
|
+
## 3 3 1
|
|
568
532
|
|
|
569
|
-
|
|
533
|
+
*dplyr* provides the ‘filter’ function, that filters data in a data
|
|
534
|
+
brame. The ‘filter’ function can be called on this data frame either by
|
|
535
|
+
using ‘R.filter(df, …)’ or by using dot notation.
|
|
570
536
|
|
|
571
|
-
|
|
572
|
-
expression. Note that if we gave to filter a Ruby expression such as
|
|
573
|
-
'x == 1', we would get an error, since there is no variable 'x' defined and if 'x' was a variable
|
|
574
|
-
then 'x == 1' would either be 'true' or 'false'. Our goal is to filter our data frame returning
|
|
575
|
-
all rows in which the 'x' value is equal to 1. To express this we want: 'R[:x].eq 1', where :x will
|
|
576
|
-
be interpreted by filter as the 'x' column.
|
|
537
|
+
——-FIX———
|
|
577
538
|
|
|
539
|
+
We prefer to use dot notation as shown below. The argument to ‘filter’
|
|
540
|
+
should be an expression. Note that if we gave to filter a Ruby
|
|
541
|
+
expression such as ‘x == 1’, we would get an error, since there is no
|
|
542
|
+
variable ‘x’ defined and if ‘x’ was a variable then ‘x == 1’ would
|
|
543
|
+
either be ‘true’ or ‘false’. Our goal is to filter our data frame
|
|
544
|
+
returning all rows in which the ‘x’ value is equal to 1. To express this
|
|
545
|
+
we want: ‘R\[:x\].eq 1’, where :x will be interpreted by filter as the
|
|
546
|
+
‘x’ column.
|
|
578
547
|
|
|
579
548
|
``` ruby
|
|
580
549
|
puts df.filter(R[:x].eq 1)
|
|
581
550
|
```
|
|
582
551
|
|
|
583
|
-
|
|
584
|
-
##
|
|
585
|
-
## 1 1 3
|
|
586
|
-
```
|
|
587
|
-
In R, and when coding with 'tidyverse', arguments to a function are usually not
|
|
588
|
-
*referencially transparent*. That is, you can’t replace a value with a seemingly equivalent
|
|
589
|
-
object that you’ve defined elsewhere. In other words, this code
|
|
552
|
+
## x y
|
|
553
|
+
## 1 1 3
|
|
590
554
|
|
|
555
|
+
In R, and when coding with ‘tidyverse’, arguments to a function are
|
|
556
|
+
usually not *referencially transparent*. That is, you can’t replace a
|
|
557
|
+
value with a seemingly equivalent object that you’ve defined elsewhere.
|
|
558
|
+
In other words, this code
|
|
591
559
|
|
|
592
560
|
``` r
|
|
593
561
|
my_var <- x
|
|
594
562
|
filter(df, my_var == 1)
|
|
595
563
|
```
|
|
596
|
-
Generates the following error: "object 'x' not found.
|
|
597
564
|
|
|
598
|
-
|
|
599
|
-
code below. Note initially that 'my_var = R[:x]' will not give the error "object 'x' not found"
|
|
600
|
-
since ':x' is treated as an expression and assigned to my\_var. Then when doing (my\_var.eq 1),
|
|
601
|
-
my\_var is a variable that resolves to ':x' and it becomes equivalent to (R[:x].eq 1) which is
|
|
602
|
-
what we want.
|
|
565
|
+
Generates the following error: “object ‘x’ not found.
|
|
603
566
|
|
|
567
|
+
However, in Galaaz, arguments are referencially transparent as can be
|
|
568
|
+
seen by the code below. Note initially that ‘my_var = R\[:x\]’ will not
|
|
569
|
+
give the error “object ‘x’ not found” since ‘:x’ is treated as an
|
|
570
|
+
expression and assigned to my_var. Then when doing (my_var.eq 1), my_var
|
|
571
|
+
is a variable that resolves to ‘:x’ and it becomes equivalent to
|
|
572
|
+
(R\[:x\].eq 1) which is what we want.
|
|
604
573
|
|
|
605
574
|
``` ruby
|
|
606
575
|
my_var = R[:x]
|
|
607
576
|
puts df.filter(my_var.eq 1)
|
|
608
577
|
```
|
|
609
578
|
|
|
610
|
-
|
|
611
|
-
##
|
|
612
|
-
|
|
613
|
-
```
|
|
579
|
+
## x y
|
|
580
|
+
## 1 1 3
|
|
581
|
+
|
|
614
582
|
As stated by Hadley
|
|
615
583
|
|
|
616
|
-
> dplyr code is ambiguous. Depending on what variables are defined
|
|
617
|
-
> filter(df, x == y) could be equivalent to any of:
|
|
584
|
+
> dplyr code is ambiguous. Depending on what variables are defined
|
|
585
|
+
> where, filter(df, x == y) could be equivalent to any of:
|
|
618
586
|
|
|
619
|
-
|
|
620
|
-
df[df$x ==
|
|
621
|
-
df[
|
|
622
|
-
df[x ==
|
|
623
|
-
df[x == y, ]
|
|
624
|
-
```
|
|
625
|
-
In galaaz this ambiguity does not exist, filter(df, x.eq y) is not a valid expression as
|
|
626
|
-
expressions are build with symbols. In doing filter(df, R[:x].eq y) we are looking for elements
|
|
627
|
-
of the 'x' column that are equal to a previously defined y variable. Finally in
|
|
628
|
-
filter(df, R[:x].eq R[:y]) we are looking for elements in which the 'x' column value is equal to
|
|
629
|
-
the 'y' column value. This can be seen in the following two chunks of code:
|
|
587
|
+
df[df$x == df$y, ]
|
|
588
|
+
df[df$x == y, ]
|
|
589
|
+
df[x == df$y, ]
|
|
590
|
+
df[x == y, ]
|
|
630
591
|
|
|
592
|
+
In galaaz this ambiguity does not exist, filter(df, x.eq y) is not a
|
|
593
|
+
valid expression as expressions are build with symbols. In doing
|
|
594
|
+
filter(df, R\[:x\].eq y) we are looking for elements of the ‘x’ column
|
|
595
|
+
that are equal to a previously defined y variable. Finally in filter(df,
|
|
596
|
+
R\[:x\].eq R\[:y\]) we are looking for elements in which the ‘x’ column
|
|
597
|
+
value is equal to the ‘y’ column value. This can be seen in the
|
|
598
|
+
following two chunks of code:
|
|
631
599
|
|
|
632
600
|
``` ruby
|
|
633
601
|
y = 1
|
|
@@ -637,11 +605,8 @@ x = 2
|
|
|
637
605
|
puts df.filter(R[:x].eq R[:y])
|
|
638
606
|
```
|
|
639
607
|
|
|
640
|
-
|
|
641
|
-
##
|
|
642
|
-
## 1 2 2
|
|
643
|
-
```
|
|
644
|
-
|
|
608
|
+
## x y
|
|
609
|
+
## 1 2 2
|
|
645
610
|
|
|
646
611
|
``` ruby
|
|
647
612
|
# looking for values where the 'x' column is equal to the 'y' variable
|
|
@@ -649,81 +614,82 @@ puts df.filter(R[:x].eq R[:y])
|
|
|
649
614
|
puts df.filter(R[:x].eq y)
|
|
650
615
|
```
|
|
651
616
|
|
|
652
|
-
|
|
653
|
-
##
|
|
654
|
-
|
|
655
|
-
```
|
|
617
|
+
## x y
|
|
618
|
+
## 1 1 3
|
|
619
|
+
|
|
656
620
|
## Writing a function that applies to different data sets
|
|
657
621
|
|
|
658
|
-
Let
|
|
659
|
-
and as second argument an expression that
|
|
660
|
-
|
|
622
|
+
Let’s suppose that we want to write a function that receives as the
|
|
623
|
+
first argument a data frame and as second argument an expression that
|
|
624
|
+
adds a column to the data frame that is equal to the sum of elements in
|
|
625
|
+
column ‘a’ plus ‘x’.
|
|
661
626
|
|
|
662
|
-
Here is the intended behaviour using the
|
|
627
|
+
Here is the intended behaviour using the ‘mutate’ function of ‘dplyr’:
|
|
628
|
+
|
|
629
|
+
mutate(df1, y = a + x)
|
|
630
|
+
mutate(df2, y = a + x)
|
|
631
|
+
mutate(df3, y = a + x)
|
|
632
|
+
mutate(df4, y = a + x)
|
|
663
633
|
|
|
664
|
-
```
|
|
665
|
-
mutate(df1, y = a + x)
|
|
666
|
-
mutate(df2, y = a + x)
|
|
667
|
-
mutate(df3, y = a + x)
|
|
668
|
-
mutate(df4, y = a + x)
|
|
669
|
-
```
|
|
670
634
|
The naive approach to writing an R function to solve this problem is:
|
|
671
635
|
|
|
672
|
-
|
|
673
|
-
|
|
674
|
-
|
|
675
|
-
}
|
|
676
|
-
```
|
|
677
|
-
Unfortunately, in R, this function can fail silently if one of the variables isn’t present
|
|
678
|
-
in the data frame, but is present in the global environment. We will not go through here how
|
|
679
|
-
to solve this problem in R.
|
|
636
|
+
mutate_y <- function(df) {
|
|
637
|
+
mutate(df, y = a + x)
|
|
638
|
+
}
|
|
680
639
|
|
|
681
|
-
|
|
640
|
+
Unfortunately, in R, this function can fail silently if one of the
|
|
641
|
+
variables isn’t present in the data frame, but is present in the global
|
|
642
|
+
environment. We will not go through here how to solve this problem in R.
|
|
682
643
|
|
|
644
|
+
In Galaaz the method mutate_y below will work fine and will never fail
|
|
645
|
+
silently.
|
|
683
646
|
|
|
684
647
|
``` ruby
|
|
685
648
|
def mutate_y(df)
|
|
686
|
-
#
|
|
649
|
+
# Column names are Ruby kwargs (y: …).
|
|
650
|
+
# Use .assign only for R `<-` expressions.
|
|
687
651
|
df.mutate(y: R[:a] + R[:x])
|
|
688
652
|
end
|
|
689
653
|
```
|
|
690
|
-
Here we create a data frame that has only one column named 'x':
|
|
691
654
|
|
|
655
|
+
Here we create a data frame that has only one column named ‘x’:
|
|
692
656
|
|
|
693
657
|
``` ruby
|
|
694
658
|
df1 = R.data__frame(x: (1..3))
|
|
695
659
|
puts df1
|
|
696
660
|
```
|
|
697
661
|
|
|
698
|
-
|
|
699
|
-
##
|
|
700
|
-
##
|
|
701
|
-
##
|
|
702
|
-
## 3 3
|
|
703
|
-
```
|
|
704
|
-
|
|
705
|
-
Note that method mutate_y will fail independetly from the fact that variable 'a' is defined and
|
|
706
|
-
in the scope of the method. Variable 'a' has no relationship with the symbol `R[:a]` used in the
|
|
707
|
-
definition of 'mutate\_y' above:
|
|
662
|
+
## x
|
|
663
|
+
## 1 1
|
|
664
|
+
## 2 2
|
|
665
|
+
## 3 3
|
|
708
666
|
|
|
667
|
+
Note that method mutate_y will fail independetly from the fact that
|
|
668
|
+
variable ‘a’ is defined and in the scope of the method. Variable ‘a’ has
|
|
669
|
+
no relationship with the symbol `R[:a]` used in the definition of
|
|
670
|
+
‘mutate_y’ above:
|
|
709
671
|
|
|
710
672
|
``` ruby
|
|
711
673
|
a = 10
|
|
712
|
-
|
|
674
|
+
begin
|
|
675
|
+
mutate_y(df1)
|
|
676
|
+
rescue => e
|
|
677
|
+
# Short message only — full backtraces overflow PDF code boxes.
|
|
678
|
+
puts "#{e.class}: #{e.message}"
|
|
679
|
+
end
|
|
713
680
|
```
|
|
714
681
|
|
|
715
|
-
|
|
716
|
-
##
|
|
717
|
-
##
|
|
718
|
-
## ! object 'a' not found
|
|
719
|
-
```
|
|
720
|
-
## Different expressions
|
|
682
|
+
## NewBridge::SessionClient::RProcessError: Error: ℹ In argument: `y = a + x`.
|
|
683
|
+
## Caused by error:
|
|
684
|
+
## ! object 'a' not found
|
|
721
685
|
|
|
722
|
-
|
|
723
|
-
that will receive two argumens, the first a variable and the second an expression is not trivial.
|
|
724
|
-
Below we create a data frame and we want to write a function that groups data by a variable and
|
|
725
|
-
summarises it by an expression:
|
|
686
|
+
## Different expressions
|
|
726
687
|
|
|
688
|
+
Let’s move to the next problem as presented by Hadley where trying to
|
|
689
|
+
write a function in R that will receive two argumens, the first a
|
|
690
|
+
variable and the second an expression is not trivial. Below we create a
|
|
691
|
+
data frame and we want to write a function that groups data by a
|
|
692
|
+
variable and summarises it by an expression:
|
|
727
693
|
|
|
728
694
|
``` r
|
|
729
695
|
set.seed(123)
|
|
@@ -738,46 +704,39 @@ df <- data.frame(
|
|
|
738
704
|
as.data.frame(df)
|
|
739
705
|
```
|
|
740
706
|
|
|
741
|
-
|
|
742
|
-
##
|
|
743
|
-
##
|
|
744
|
-
## 2 1
|
|
745
|
-
##
|
|
746
|
-
##
|
|
747
|
-
## 5 2 1 1 4
|
|
748
|
-
```
|
|
707
|
+
## g1 g2 a b
|
|
708
|
+
## 1 1 1 3 3
|
|
709
|
+
## 2 1 2 2 1
|
|
710
|
+
## 3 2 1 5 2
|
|
711
|
+
## 4 2 2 4 5
|
|
712
|
+
## 5 2 1 1 4
|
|
749
713
|
|
|
750
714
|
``` r
|
|
751
715
|
d2 <- df %>%
|
|
752
716
|
group_by(g1) %>%
|
|
753
717
|
summarise(a = mean(a))
|
|
754
|
-
|
|
718
|
+
|
|
755
719
|
as.data.frame(d2)
|
|
756
720
|
```
|
|
757
721
|
|
|
758
|
-
|
|
759
|
-
##
|
|
760
|
-
##
|
|
761
|
-
## 2 2 3.333333
|
|
762
|
-
```
|
|
722
|
+
## g1 a
|
|
723
|
+
## 1 1 2.500000
|
|
724
|
+
## 2 2 3.333333
|
|
763
725
|
|
|
764
726
|
``` r
|
|
765
727
|
d2 <- df %>%
|
|
766
728
|
group_by(g2) %>%
|
|
767
729
|
summarise(a = mean(a))
|
|
768
|
-
|
|
769
|
-
as.data.frame(d2)
|
|
730
|
+
|
|
731
|
+
as.data.frame(d2)
|
|
770
732
|
```
|
|
771
733
|
|
|
772
|
-
|
|
773
|
-
##
|
|
774
|
-
##
|
|
775
|
-
## 2 2 3
|
|
776
|
-
```
|
|
734
|
+
## g2 a
|
|
735
|
+
## 1 1 3
|
|
736
|
+
## 2 2 3
|
|
777
737
|
|
|
778
738
|
As shown by Hadley, one might expect this function to do the trick:
|
|
779
739
|
|
|
780
|
-
|
|
781
740
|
``` r
|
|
782
741
|
my_summarise <- function(df, group_var) {
|
|
783
742
|
df %>%
|
|
@@ -789,15 +748,17 @@ my_summarise <- function(df, group_var) {
|
|
|
789
748
|
#> Error: Column `group_var` is unknown
|
|
790
749
|
```
|
|
791
750
|
|
|
792
|
-
In order to solve this problem, coding with dplyr requires the
|
|
793
|
-
and functions such as
|
|
794
|
-
|
|
795
|
-
|
|
796
|
-
Now, let's try to implement the same function in galaaz. The next code block first prints the
|
|
797
|
-
'df' data frame defined previously in R (to access an R variable from Galaaz, we use the tilde
|
|
798
|
-
operator '~' applied to the R variable name as symbol, i.e., ':df'. We then create the
|
|
799
|
-
'my_summarize' method and call it passing the R data frame and the group by variable ':g1':
|
|
751
|
+
In order to solve this problem, coding with dplyr requires the
|
|
752
|
+
introduction of many new concepts and functions such as ‘quo’, ‘quos’,
|
|
753
|
+
‘enquo’, ‘enquos’, ‘!!’ (bang bang), ‘!!!’ (triple bang). Again, we’ll
|
|
754
|
+
leave to Hadley the explanation on how to use all those functions.
|
|
800
755
|
|
|
756
|
+
Now, let’s try to implement the same function in galaaz. The next code
|
|
757
|
+
block first prints the ‘df’ data frame defined previously in R (to
|
|
758
|
+
access an R variable from Galaaz, we use the tilde operator ‘~’ applied
|
|
759
|
+
to the R variable name as symbol, i.e., ‘:df’. We then create the
|
|
760
|
+
‘my_summarize’ method and call it passing the R data frame and the group
|
|
761
|
+
by variable ‘:g1’:
|
|
801
762
|
|
|
802
763
|
``` ruby
|
|
803
764
|
puts ~R[:df]
|
|
@@ -812,64 +773,59 @@ end
|
|
|
812
773
|
puts my_summarize(~R[:df], R[:g1])
|
|
813
774
|
```
|
|
814
775
|
|
|
815
|
-
|
|
816
|
-
##
|
|
817
|
-
##
|
|
818
|
-
## 2 1
|
|
819
|
-
##
|
|
820
|
-
##
|
|
821
|
-
##
|
|
822
|
-
##
|
|
823
|
-
##
|
|
824
|
-
##
|
|
825
|
-
##
|
|
826
|
-
##
|
|
827
|
-
## 2 2 3.33
|
|
828
|
-
```
|
|
829
|
-
It works!!! Well, let's make sure this was not just some coincidence
|
|
776
|
+
## g1 g2 a b
|
|
777
|
+
## 1 1 1 3 3
|
|
778
|
+
## 2 1 2 2 1
|
|
779
|
+
## 3 2 1 5 2
|
|
780
|
+
## 4 2 2 4 5
|
|
781
|
+
## 5 2 1 1 4
|
|
782
|
+
##
|
|
783
|
+
## # A tibble: 2 × 2
|
|
784
|
+
## g1 a
|
|
785
|
+
## <dbl> <dbl>
|
|
786
|
+
## 1 1 2.5
|
|
787
|
+
## 2 2 3.33
|
|
830
788
|
|
|
789
|
+
It works!!! Well, let’s make sure this was not just some coincidence
|
|
831
790
|
|
|
832
791
|
``` ruby
|
|
833
792
|
puts my_summarize(~R[:df], R[:g2])
|
|
834
793
|
```
|
|
835
794
|
|
|
836
|
-
|
|
837
|
-
##
|
|
838
|
-
##
|
|
839
|
-
##
|
|
840
|
-
##
|
|
841
|
-
## 2 2 3
|
|
842
|
-
```
|
|
795
|
+
## # A tibble: 2 × 2
|
|
796
|
+
## g2 a
|
|
797
|
+
## <dbl> <dbl>
|
|
798
|
+
## 1 1 3
|
|
799
|
+
## 2 2 3
|
|
843
800
|
|
|
844
|
-
Great, everything is fine! No magic, no new functions, no complexities,
|
|
845
|
-
code.
|
|
801
|
+
Great, everything is fine! No magic, no new functions, no complexities,
|
|
802
|
+
just normal, standard Ruby code. If you’ve ever done NSE in R, this
|
|
803
|
+
certainly feels much safer and easy to implement.
|
|
846
804
|
|
|
847
805
|
## Different input variables
|
|
848
806
|
|
|
849
|
-
In the previous section we
|
|
850
|
-
does this remain true for more complex
|
|
851
|
-
more complex
|
|
807
|
+
In the previous section we’ve managed to get rid of all NSE formulation
|
|
808
|
+
for a simple example, but does this remain true for more complex
|
|
809
|
+
examples, or will the Galaaz way prove inpractical for more complex
|
|
810
|
+
code?
|
|
852
811
|
|
|
853
|
-
In the next example Hadley proposes us to write a function that given an
|
|
854
|
-
or
|
|
855
|
-
statements:
|
|
812
|
+
In the next example Hadley proposes us to write a function that given an
|
|
813
|
+
expression such as ‘a’ or ‘a \* b’, calculates three summaries. What we
|
|
814
|
+
want a function that does the same as these R statements:
|
|
856
815
|
|
|
857
|
-
|
|
858
|
-
|
|
859
|
-
#>
|
|
860
|
-
#>
|
|
861
|
-
#>
|
|
862
|
-
#> 1 3 15 5
|
|
863
|
-
|
|
864
|
-
summarise(df, mean = mean(a * b), sum = sum(a * b), n = n())
|
|
865
|
-
#> # A tibble: 1 x 3
|
|
866
|
-
#> mean sum n
|
|
867
|
-
#> <dbl> <int> <int>
|
|
868
|
-
#> 1 9 45 5
|
|
869
|
-
```
|
|
816
|
+
summarise(df, mean = mean(a), sum = sum(a), n = n())
|
|
817
|
+
#> # A tibble: 1 x 3
|
|
818
|
+
#> mean sum n
|
|
819
|
+
#> <dbl> <int> <int>
|
|
820
|
+
#> 1 3 15 5
|
|
870
821
|
|
|
871
|
-
|
|
822
|
+
summarise(df, mean = mean(a * b), sum = sum(a * b), n = n())
|
|
823
|
+
#> # A tibble: 1 x 3
|
|
824
|
+
#> mean sum n
|
|
825
|
+
#> <dbl> <int> <int>
|
|
826
|
+
#> 1 9 45 5
|
|
872
827
|
|
|
828
|
+
Let’s try it in galaaz:
|
|
873
829
|
|
|
874
830
|
``` ruby
|
|
875
831
|
def my_summarise2(df, expr)
|
|
@@ -884,50 +840,49 @@ puts my_summarise2((~R[:df]), :a)
|
|
|
884
840
|
puts my_summarise2((~R[:df]), R[:a] * R[:b])
|
|
885
841
|
```
|
|
886
842
|
|
|
887
|
-
|
|
888
|
-
##
|
|
889
|
-
##
|
|
890
|
-
##
|
|
891
|
-
## 1 9 45 5
|
|
892
|
-
```
|
|
843
|
+
## mean sum n
|
|
844
|
+
## 1 3 15 5
|
|
845
|
+
## mean sum n
|
|
846
|
+
## 1 9 45 5
|
|
893
847
|
|
|
894
|
-
Once again, there is no need to use any special theory or functions.
|
|
895
|
-
careful about is the use of
|
|
848
|
+
Once again, there is no need to use any special theory or functions. The
|
|
849
|
+
only point to be careful about is the use of ‘E’ to build expressions
|
|
850
|
+
from functions ‘mean’, ‘sum’ and ‘n’.
|
|
896
851
|
|
|
897
852
|
## Different input and output variable
|
|
898
853
|
|
|
899
|
-
Now the next challenge presented by Hadley is to vary the name of the
|
|
900
|
-
the received expression.
|
|
901
|
-
|
|
902
|
-
|
|
903
|
-
|
|
904
|
-
|
|
905
|
-
mutate(df, mean_a = mean(a), sum_a = sum(a))
|
|
906
|
-
#> # A tibble: 5 x 6
|
|
907
|
-
#> g1 g2 a b mean_a sum_a
|
|
908
|
-
#> <dbl> <dbl> <int> <int> <dbl> <int>
|
|
909
|
-
#> 1 1 1 1 3 3 15
|
|
910
|
-
#> 2 1 2 4 2 3 15
|
|
911
|
-
#> 3 2 1 2 1 3 15
|
|
912
|
-
#> 4 2 2 5 4 3 15
|
|
913
|
-
#> # … with 1 more row
|
|
914
|
-
|
|
915
|
-
mutate(df, mean_b = mean(b), sum_b = sum(b))
|
|
916
|
-
#> # A tibble: 5 x 6
|
|
917
|
-
#> g1 g2 a b mean_b sum_b
|
|
918
|
-
#> <dbl> <dbl> <int> <int> <dbl> <int>
|
|
919
|
-
#> 1 1 1 1 3 3 15
|
|
920
|
-
#> 2 1 2 4 2 3 15
|
|
921
|
-
#> 3 2 1 2 1 3 15
|
|
922
|
-
#> 4 2 2 5 4 3 15
|
|
923
|
-
#> # … with 1 more row
|
|
924
|
-
|
|
925
|
-
In order to solve this problem in R, Hadley needs to introduce some more
|
|
926
|
-
|
|
854
|
+
Now the next challenge presented by Hadley is to vary the name of the
|
|
855
|
+
output variables based on the received expression. So, if the input
|
|
856
|
+
expression is ‘a’, we want our data frame columns to be named ‘mean_a’
|
|
857
|
+
and ‘sum_a’. Now, if the input expression is ‘b’, columns should be
|
|
858
|
+
named ‘mean_b’ and ‘sum_b’.
|
|
859
|
+
|
|
860
|
+
mutate(df, mean_a = mean(a), sum_a = sum(a))
|
|
861
|
+
#> # A tibble: 5 x 6
|
|
862
|
+
#> g1 g2 a b mean_a sum_a
|
|
863
|
+
#> <dbl> <dbl> <int> <int> <dbl> <int>
|
|
864
|
+
#> 1 1 1 1 3 3 15
|
|
865
|
+
#> 2 1 2 4 2 3 15
|
|
866
|
+
#> 3 2 1 2 1 3 15
|
|
867
|
+
#> 4 2 2 5 4 3 15
|
|
868
|
+
#> # … with 1 more row
|
|
869
|
+
|
|
870
|
+
mutate(df, mean_b = mean(b), sum_b = sum(b))
|
|
871
|
+
#> # A tibble: 5 x 6
|
|
872
|
+
#> g1 g2 a b mean_b sum_b
|
|
873
|
+
#> <dbl> <dbl> <int> <int> <dbl> <int>
|
|
874
|
+
#> 1 1 1 1 3 3 15
|
|
875
|
+
#> 2 1 2 4 2 3 15
|
|
876
|
+
#> 3 2 1 2 1 3 15
|
|
877
|
+
#> 4 2 2 5 4 3 15
|
|
878
|
+
#> # … with 1 more row
|
|
879
|
+
|
|
880
|
+
In order to solve this problem in R, Hadley needs to introduce some more
|
|
881
|
+
new functions and notations: ‘quo_name’ and the ‘:=’ operator from
|
|
882
|
+
package ‘rlang’
|
|
927
883
|
|
|
928
884
|
Here is our Ruby code:
|
|
929
885
|
|
|
930
|
-
|
|
931
886
|
``` ruby
|
|
932
887
|
def my_mutate(df, expr)
|
|
933
888
|
mean_name = "mean_#{expr.to_s}"
|
|
@@ -941,36 +896,37 @@ puts my_mutate((~R[:df]), :a)
|
|
|
941
896
|
puts my_mutate((~R[:df]), :b)
|
|
942
897
|
```
|
|
943
898
|
|
|
944
|
-
|
|
945
|
-
##
|
|
946
|
-
##
|
|
947
|
-
## 2 1
|
|
948
|
-
##
|
|
949
|
-
##
|
|
950
|
-
##
|
|
951
|
-
##
|
|
952
|
-
##
|
|
953
|
-
## 2 1
|
|
954
|
-
##
|
|
955
|
-
##
|
|
956
|
-
|
|
957
|
-
|
|
958
|
-
|
|
959
|
-
|
|
960
|
-
|
|
961
|
-
followed by a
|
|
962
|
-
|
|
963
|
-
|
|
964
|
-
|
|
899
|
+
## g1 g2 a b mean_a sum_a
|
|
900
|
+
## 1 1 1 3 3 3 15
|
|
901
|
+
## 2 1 2 2 1 3 15
|
|
902
|
+
## 3 2 1 5 2 3 15
|
|
903
|
+
## 4 2 2 4 5 3 15
|
|
904
|
+
## 5 2 1 1 4 3 15
|
|
905
|
+
## g1 g2 a b mean_b sum_b
|
|
906
|
+
## 1 1 1 3 3 3 15
|
|
907
|
+
## 2 1 2 2 1 3 15
|
|
908
|
+
## 3 2 1 5 2 3 15
|
|
909
|
+
## 4 2 2 4 5 3 15
|
|
910
|
+
## 5 2 1 1 4 3 15
|
|
911
|
+
|
|
912
|
+
It really seems that “Non Standard Evaluation” is actually quite
|
|
913
|
+
standard in Galaaz! But, you might have noticed a small change in the
|
|
914
|
+
way the arguments to the mutate method were called. In a previous
|
|
915
|
+
example we used df.summarise(mean: E.mean(:a), …) where the column name
|
|
916
|
+
was followed by a ‘:’ colom. In this example, we have
|
|
917
|
+
df.mutate(mean_name =\> E.mean(expr), …) and variable mean_name is not
|
|
918
|
+
followed by ‘:’ but by ‘=\>’. This is standard Ruby notation.
|
|
919
|
+
|
|
920
|
+
\[explain….\]
|
|
965
921
|
|
|
966
922
|
## Capturing multiple variables
|
|
967
923
|
|
|
968
|
-
Moving on with new complexities, Hadley proposes us to solve the problem
|
|
969
|
-
summarise function will receive any number of grouping
|
|
970
|
-
|
|
971
|
-
This again is quite standard Ruby. In order to receive an undefined number of paramenters
|
|
972
|
-
the paramenter is preceded by '*':
|
|
924
|
+
Moving on with new complexities, Hadley proposes us to solve the problem
|
|
925
|
+
in which the summarise function will receive any number of grouping
|
|
926
|
+
variables.
|
|
973
927
|
|
|
928
|
+
This again is quite standard Ruby. In order to receive an undefined
|
|
929
|
+
number of paramenters the paramenter is preceded by ’\*’:
|
|
974
930
|
|
|
975
931
|
``` ruby
|
|
976
932
|
def my_summarise3(df, *group_vars)
|
|
@@ -981,85 +937,89 @@ end
|
|
|
981
937
|
puts my_summarise3((~R[:df]), :g1, :g2)
|
|
982
938
|
```
|
|
983
939
|
|
|
984
|
-
|
|
985
|
-
## #
|
|
986
|
-
##
|
|
987
|
-
##
|
|
988
|
-
##
|
|
989
|
-
##
|
|
990
|
-
## 2 1
|
|
991
|
-
##
|
|
992
|
-
## 4 2 2 4
|
|
993
|
-
```
|
|
940
|
+
## # A tibble: 4 × 3
|
|
941
|
+
## # Groups: g1 [2]
|
|
942
|
+
## g1 g2 a
|
|
943
|
+
## <dbl> <dbl> <dbl>
|
|
944
|
+
## 1 1 1 3
|
|
945
|
+
## 2 1 2 2
|
|
946
|
+
## 3 2 1 3
|
|
947
|
+
## 4 2 2 4
|
|
994
948
|
|
|
995
949
|
# Why does R require NSE and Galaaz does not?
|
|
996
950
|
|
|
997
|
-
NSE introduces a number of new concepts, such as
|
|
998
|
-
|
|
999
|
-
|
|
1000
|
-
|
|
1001
|
-
|
|
1002
|
-
|
|
1003
|
-
|
|
1004
|
-
|
|
1005
|
-
|
|
1006
|
-
|
|
1007
|
-
|
|
1008
|
-
|
|
1009
|
-
|
|
1010
|
-
|
|
1011
|
-
|
|
1012
|
-
|
|
1013
|
-
|
|
1014
|
-
|
|
1015
|
-
|
|
1016
|
-
|
|
1017
|
-
|
|
1018
|
-
|
|
1019
|
-
|
|
1020
|
-
|
|
1021
|
-
|
|
951
|
+
NSE introduces a number of new concepts, such as ‘quoting’,
|
|
952
|
+
‘quasiquotation’, ‘unquoting’ and ‘unquote-splicing’, while in Galaaz
|
|
953
|
+
none of those concepts are needed. What gives?
|
|
954
|
+
|
|
955
|
+
R is an extremely flexible language and it has lazy evaluation of
|
|
956
|
+
parameters. When in R a function is called as ‘summarise(df, a = b)’,
|
|
957
|
+
the summarise function receives the litteral ‘a = b’ parameter and can
|
|
958
|
+
work with this as if it were a string. In R, it is not clear what a and
|
|
959
|
+
b are, they can be expressions or they can be variables, it is up to the
|
|
960
|
+
function to decide what ‘a = b’ means.
|
|
961
|
+
|
|
962
|
+
In Ruby, there is no lazy evaluation of parameters and ‘a’ is always a
|
|
963
|
+
variable and so is ‘b’. Variables assume their value as soon as they are
|
|
964
|
+
used, so ‘x = a’ is immediately evaluate and variable ‘x’ will receive
|
|
965
|
+
the value of variable ‘a’ as soon as the Ruby statement is executed.
|
|
966
|
+
Ruby also provides the notion of a symbol; ‘:a’ is a symbol and does not
|
|
967
|
+
evaluate to anything. Galaaz uses Ruby symbols to build expressions that
|
|
968
|
+
are not bound to anything: ‘R\[:a\].eq R\[:b\]’ is clearly an expression
|
|
969
|
+
and has no relationship whatsoever with the statment ‘a = b’. By using
|
|
970
|
+
symbols, variables and expressions all the possible ambiguities that are
|
|
971
|
+
found in R are eliminated in Galaaz.
|
|
972
|
+
|
|
973
|
+
The main problem that remains, is that in R, functions are not clearly
|
|
974
|
+
documented as what type of input they are expecting, they might be
|
|
975
|
+
expecting regular variables or they might be expecting expressions and
|
|
976
|
+
the R function will know how to deal with an input of the form ‘a = b’,
|
|
977
|
+
now for the Ruby developer it might not be immediately clear if it
|
|
978
|
+
should call the function passing the value ‘true’ if variable ‘a’ is
|
|
979
|
+
equal to variable ‘b’ or if it should call the function passing the
|
|
980
|
+
expression ‘R\[:a\].eq R\[:b\]’.
|
|
1022
981
|
|
|
1023
982
|
# Advanced dplyr features
|
|
1024
983
|
|
|
1025
|
-
In the blog: [Programming with dplyr by using
|
|
1026
|
-
|
|
1027
|
-
|
|
1028
|
-
|
|
1029
|
-
> program over dplyr without having “to bring in (or study) any deep-theory or
|
|
1030
|
-
> heavy-weight tools such as rlang/tidyeval”.
|
|
984
|
+
In the blog: [Programming with dplyr by using
|
|
985
|
+
dplyr](https://www.r-bloggers.com/programming-with-dplyr-by-using-dplyr/)
|
|
986
|
+
Iñaki Úcar shows surprise that some R users are trying to code in dplyr
|
|
987
|
+
avoiding the use of NSE. For instance he says:
|
|
1031
988
|
|
|
1032
|
-
|
|
1033
|
-
|
|
1034
|
-
|
|
1035
|
-
a 'piece of cake'. So much so, that 'tidyeval' has some more advanced functions that instead
|
|
1036
|
-
of using quoted expressions, uses strings as arguments.
|
|
989
|
+
> Take the example of seplyr. It stands for standard evaluation dplyr,
|
|
990
|
+
> and enables us to program over dplyr without having “to bring in (or
|
|
991
|
+
> study) any deep-theory or heavy-weight tools such as rlang/tidyeval”.
|
|
1037
992
|
|
|
1038
|
-
|
|
1039
|
-
|
|
1040
|
-
|
|
993
|
+
For me, there isn’t really any surprise that users are trying to avoid
|
|
994
|
+
dplyr deep-theory. R users frequently are not programmers and learning
|
|
995
|
+
to code is already hard business, on top of that, having to learn how to
|
|
996
|
+
‘quote’ or ‘enquo’ or ‘quos’ or ‘enquos’ is not necessarily a ‘piece of
|
|
997
|
+
cake’. So much so, that ‘tidyeval’ has some more advanced functions that
|
|
998
|
+
instead of using quoted expressions, uses strings as arguments.
|
|
1041
999
|
|
|
1000
|
+
In the following examples, we show the use of functions ‘group_by_at’,
|
|
1001
|
+
‘summarise_at’ and ‘rename_at’ that receive strings as argument. The
|
|
1002
|
+
data frame used in ‘starwars’ that describes features of characters in
|
|
1003
|
+
the Starwars movies:
|
|
1042
1004
|
|
|
1043
1005
|
``` ruby
|
|
1044
1006
|
puts (~R[:starwars]).head
|
|
1045
1007
|
```
|
|
1046
1008
|
|
|
1047
|
-
|
|
1048
|
-
##
|
|
1049
|
-
##
|
|
1050
|
-
##
|
|
1051
|
-
##
|
|
1052
|
-
##
|
|
1053
|
-
##
|
|
1054
|
-
##
|
|
1055
|
-
##
|
|
1056
|
-
## 6
|
|
1057
|
-
## #
|
|
1058
|
-
## # vehicles <list>, starships <list>
|
|
1059
|
-
```
|
|
1060
|
-
The grouped_mean function below will receive a grouping variable and calculate summaries for
|
|
1061
|
-
the value\_variables given:
|
|
1009
|
+
## # A tibble: 6 × 14
|
|
1010
|
+
## name height mass hair_color skin_color eye_color birth_year sex
|
|
1011
|
+
## <chr> <int> <dbl> <chr> <chr> <chr> <dbl> <chr>
|
|
1012
|
+
## 1 Luke … 172 77 blond fair blue 19 male
|
|
1013
|
+
## 2 C-3PO 167 75 <NA> gold yellow 112 none
|
|
1014
|
+
## 3 R2-D2 96 32 <NA> white, bl… red 33 none
|
|
1015
|
+
## 4 Darth… 202 136 none white yellow 41.9 male
|
|
1016
|
+
## 5 Leia … 150 49 brown light brown 19 fema…
|
|
1017
|
+
## 6 Owen … 178 120 brown, gr… light blue 52 male
|
|
1018
|
+
## # ℹ 6 more variables: gender <chr>, homeworld <chr>, species <chr>,
|
|
1019
|
+
## # films <list>, vehicles <list>, starships <list>
|
|
1062
1020
|
|
|
1021
|
+
The grouped_mean function below will receive a grouping variable and
|
|
1022
|
+
calculate summaries for the value_variables given:
|
|
1063
1023
|
|
|
1064
1024
|
``` r
|
|
1065
1025
|
grouped_mean <- function(data, grouping_variables, value_variables) {
|
|
@@ -1074,94 +1034,105 @@ gm = starwars %>%
|
|
|
1074
1034
|
grouped_mean("eye_color", c("mass", "birth_year"))
|
|
1075
1035
|
```
|
|
1076
1036
|
|
|
1077
|
-
|
|
1078
|
-
##
|
|
1079
|
-
##
|
|
1080
|
-
##
|
|
1081
|
-
##
|
|
1082
|
-
##
|
|
1083
|
-
##
|
|
1084
|
-
##
|
|
1085
|
-
##
|
|
1086
|
-
##
|
|
1087
|
-
## generated.
|
|
1088
|
-
```
|
|
1037
|
+
## Warning: `funs()` was deprecated in dplyr 0.8.0.
|
|
1038
|
+
## ℹ Please use a list of either functions or lambdas:
|
|
1039
|
+
##
|
|
1040
|
+
## # Simple named list: list(mean = mean, median = median)
|
|
1041
|
+
##
|
|
1042
|
+
## # Auto named with `tibble::lst()`: tibble::lst(mean, median)
|
|
1043
|
+
##
|
|
1044
|
+
## # Using lambdas list(~ mean(., trim = .2), ~ median(., na.rm = TRUE))
|
|
1045
|
+
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning
|
|
1046
|
+
## was generated.
|
|
1089
1047
|
|
|
1090
1048
|
``` r
|
|
1091
1049
|
as.data.frame(gm)
|
|
1092
1050
|
```
|
|
1093
1051
|
|
|
1094
|
-
|
|
1095
|
-
##
|
|
1096
|
-
##
|
|
1097
|
-
##
|
|
1098
|
-
##
|
|
1099
|
-
##
|
|
1100
|
-
##
|
|
1101
|
-
##
|
|
1102
|
-
##
|
|
1103
|
-
##
|
|
1104
|
-
##
|
|
1105
|
-
##
|
|
1106
|
-
##
|
|
1107
|
-
##
|
|
1108
|
-
##
|
|
1109
|
-
##
|
|
1110
|
-
## 15 yellow 81.11111 76.38000 11
|
|
1111
|
-
```
|
|
1052
|
+
## eye_color mean_mass mean_birth_year count
|
|
1053
|
+
## 1 black 76.28571 33.00000 10
|
|
1054
|
+
## 2 blue 86.51667 67.06923 19
|
|
1055
|
+
## 3 blue-gray 77.00000 57.00000 1
|
|
1056
|
+
## 4 brown 66.09231 108.96429 21
|
|
1057
|
+
## 5 dark NaN NaN 1
|
|
1058
|
+
## 6 gold NaN NaN 1
|
|
1059
|
+
## 7 green, yellow 159.00000 NaN 1
|
|
1060
|
+
## 8 hazel 66.00000 34.50000 3
|
|
1061
|
+
## 9 orange 282.33333 231.00000 8
|
|
1062
|
+
## 10 pink NaN NaN 1
|
|
1063
|
+
## 11 red 81.40000 33.66667 5
|
|
1064
|
+
## 12 red, blue NaN NaN 1
|
|
1065
|
+
## 13 unknown 31.50000 NaN 3
|
|
1066
|
+
## 14 white 48.00000 NaN 1
|
|
1067
|
+
## 15 yellow 81.11111 76.38000 11
|
|
1112
1068
|
|
|
1113
1069
|
The same code with Galaaz, becomes:
|
|
1114
1070
|
|
|
1115
|
-
|
|
1116
1071
|
``` ruby
|
|
1117
1072
|
def grouped_mean(data, grouping_variables, value_variables)
|
|
1118
1073
|
data.
|
|
1119
1074
|
group_by_at(grouping_variables).
|
|
1120
1075
|
mutate(count: E.n).
|
|
1121
|
-
summarise_at(
|
|
1122
|
-
|
|
1076
|
+
summarise_at(
|
|
1077
|
+
E.c(value_variables, "count"),
|
|
1078
|
+
R[:mean],
|
|
1079
|
+
na__rm: true).
|
|
1080
|
+
rename_at(
|
|
1081
|
+
value_variables,
|
|
1082
|
+
E.funs(E.paste0("mean_", value_variables)))
|
|
1123
1083
|
end
|
|
1124
1084
|
|
|
1125
|
-
puts grouped_mean(
|
|
1126
|
-
|
|
1127
|
-
|
|
1128
|
-
|
|
1129
|
-
|
|
1130
|
-
|
|
1131
|
-
##
|
|
1132
|
-
##
|
|
1133
|
-
##
|
|
1134
|
-
##
|
|
1135
|
-
##
|
|
1136
|
-
##
|
|
1137
|
-
##
|
|
1138
|
-
##
|
|
1139
|
-
##
|
|
1140
|
-
##
|
|
1141
|
-
##
|
|
1142
|
-
##
|
|
1143
|
-
##
|
|
1144
|
-
##
|
|
1145
|
-
##
|
|
1146
|
-
##
|
|
1147
|
-
|
|
1085
|
+
puts grouped_mean(
|
|
1086
|
+
(~R[:starwars]),
|
|
1087
|
+
"eye_color",
|
|
1088
|
+
E.c("mass", "birth_year"))
|
|
1089
|
+
```
|
|
1090
|
+
|
|
1091
|
+
## # A tibble: 15 × 4
|
|
1092
|
+
## eye_color mean_mass mean_birth_year count
|
|
1093
|
+
## <chr> <dbl> <dbl> <dbl>
|
|
1094
|
+
## 1 black 76.3 33 10
|
|
1095
|
+
## 2 blue 86.5 67.1 19
|
|
1096
|
+
## 3 blue-gray 77 57 1
|
|
1097
|
+
## 4 brown 66.1 109. 21
|
|
1098
|
+
## 5 dark NaN NaN 1
|
|
1099
|
+
## 6 gold NaN NaN 1
|
|
1100
|
+
## 7 green, yellow 159 NaN 1
|
|
1101
|
+
## 8 hazel 66 34.5 3
|
|
1102
|
+
## 9 orange 282. 231 8
|
|
1103
|
+
## 10 pink NaN NaN 1
|
|
1104
|
+
## 11 red 81.4 33.7 5
|
|
1105
|
+
## 12 red, blue NaN NaN 1
|
|
1106
|
+
## 13 unknown 31.5 NaN 3
|
|
1107
|
+
## 14 white 48 NaN 1
|
|
1108
|
+
## 15 yellow 81.1 76.4 11
|
|
1148
1109
|
|
|
1149
1110
|
# Further reading
|
|
1150
1111
|
|
|
1151
|
-
|
|
1152
|
-
|
|
1153
|
-
|
|
1154
|
-
|
|
1155
|
-
|
|
1156
|
-
|
|
1157
|
-
|
|
1112
|
+
- [JRuby](https://www.jruby.org/) — Ruby on the JVM (Galaaz 2.0)
|
|
1113
|
+
- [How to make Beautiful Ruby Plots with
|
|
1114
|
+
Galaaz](https://medium.freecodecamp.org/how-to-make-beautiful-ruby-plots-with-galaaz-320848058857)
|
|
1115
|
+
(plots; narrative partly pre-2.0)
|
|
1116
|
+
- [Ruby Plotting with Galaaz in
|
|
1117
|
+
GraalVM](https://towardsdatascience.com/ruby-plotting-with-galaaz-an-example-of-tightly-coupling-ruby-and-r-in-graalvm-520b69e21021)
|
|
1118
|
+
(older stack; ideas still useful)
|
|
1119
|
+
- [How to do reproducible research in Ruby with
|
|
1120
|
+
gKnit](https://towardsdatascience.com/how-to-do-reproducible-research-in-ruby-with-gknit-c26d2684d64e)
|
|
1121
|
+
- [R for Data Science](https://r4ds.had.co.nz/)
|
|
1122
|
+
- [Advanced R](https://adv-r.hadley.nz/)
|
|
1123
|
+
- Historical context: [GraalVM](https://www.graalvm.org/),
|
|
1124
|
+
[TruffleRuby](https://github.com/oracle/truffleruby),
|
|
1125
|
+
[FastR](https://github.com/oracle/fastr)
|
|
1158
1126
|
|
|
1159
1127
|
# Conclusion
|
|
1160
1128
|
|
|
1161
|
-
Ruby and Galaaz provide a nice framework for developing code that uses R
|
|
1162
|
-
a very powerful and flexible language,
|
|
1163
|
-
|
|
1164
|
-
|
|
1165
|
-
|
|
1166
|
-
|
|
1167
|
-
|
|
1129
|
+
Ruby and Galaaz provide a nice framework for developing code that uses R
|
|
1130
|
+
functions. Although R is a very powerful and flexible language,
|
|
1131
|
+
sometimes, too much flexibility makes life harder for the casual user.
|
|
1132
|
+
We believe however, that even for the advanced user, Ruby integrated
|
|
1133
|
+
with R throught Galaaz, makes a powerful environment for data analysis.
|
|
1134
|
+
In this blog post we showed how Galaaz consistent syntax eliminates the
|
|
1135
|
+
need for complex constructs such as quoting, enquoting, quasiquotation,
|
|
1136
|
+
etc. This simplification comes from the fact that expressions and
|
|
1137
|
+
variables are clearly separated objects, which is not the case in the R
|
|
1138
|
+
language.
|