galaaz 2.1.8 → 2.1.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +30 -0
- data/Rakefile +20 -2
- data/bin/check_gemfile_lock_version +46 -0
- data/bin/release_bump +26 -0
- data/blogs/README.md +4 -0
- data/blogs/galaaz_2_0/galaaz_2_0.Rmd +385 -0
- data/blogs/galaaz_2_0/galaaz_2_0.md +409 -0
- data/blogs/galaaz_2_0/galaaz_2_0.tex +756 -0
- data/blogs/galaaz_2_0/images/galaaz-header.png +0 -0
- data/blogs/galaaz_2_0/images/galaaz-lockup-stacked.png +0 -0
- data/blogs/galaaz_ggplot/galaaz_ggplot.Rmd +14 -1
- data/blogs/galaaz_ggplot/galaaz_ggplot.md +123 -103
- data/blogs/galaaz_ggplot/galaaz_ggplot.tex +60 -23
- data/blogs/galaaz_ggplot/images/galaaz-lockup-stacked.png +0 -0
- data/blogs/gknit/gknit.Rmd +16 -1
- data/blogs/gknit/gknit.md +13 -1
- data/blogs/gknit/gknit.tex +64 -23
- data/blogs/gknit/gknit_files/figure-html/bubble-1.png +0 -0
- data/blogs/gknit/gknit_files/figure-html/diverging_bar.png +0 -0
- data/blogs/gknit/gknit_files/figure-latex/bubble-1.png +0 -0
- data/blogs/gknit/images/galaaz-lockup-stacked.png +0 -0
- data/blogs/manual/images/galaaz-lockup-stacked.png +0 -0
- data/blogs/manual/manual.Rmd +32 -11
- data/blogs/manual/manual.md +30 -23
- data/blogs/manual/manual.tex +88 -66
- data/blogs/manual/manual_files/figure-html/bubble-1.png +0 -0
- data/blogs/manual/manual_files/figure-latex/bubble-1.png +0 -0
- data/blogs/nse_dplyr/images/galaaz-lockup-stacked.png +0 -0
- data/blogs/nse_dplyr/nse_dplyr.Rmd +14 -1
- data/blogs/nse_dplyr/nse_dplyr.md +697 -649
- data/blogs/nse_dplyr/nse_dplyr.tex +61 -24
- data/blogs/oh_my/images/galaaz-lockup-stacked.png +0 -0
- data/blogs/oh_my/oh_my.Rmd +14 -1
- data/blogs/oh_my/oh_my.md +36 -26
- data/blogs/oh_my/oh_my.tex +95 -58
- data/blogs/r_on_rails_ledger/images/00_portfolio_page.png +0 -0
- data/blogs/r_on_rails_ledger/images/01_results_panel.png +0 -0
- data/blogs/r_on_rails_ledger/images/02_density_tail_risk.png +0 -0
- data/blogs/r_on_rails_ledger/images/03_mc_cone.png +0 -0
- data/blogs/r_on_rails_ledger/images/04_rolling_var.png +0 -0
- data/blogs/r_on_rails_ledger/images/galaaz-lockup-stacked.png +0 -0
- data/blogs/r_on_rails_ledger/r_on_rails_ledger.Rmd +354 -0
- data/blogs/r_on_rails_ledger/r_on_rails_ledger.md +365 -0
- data/blogs/r_on_rails_ledger/r_on_rails_ledger.tex +670 -0
- data/blogs/ruby_plot/images/galaaz-lockup-stacked.png +0 -0
- data/blogs/ruby_plot/ruby_plot.Rmd +14 -1
- data/blogs/ruby_plot/ruby_plot.md +11 -1
- data/blogs/ruby_plot/ruby_plot.tex +60 -23
- data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/final_violin_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-html/violin_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/final_violin_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/figure-latex/violin_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_with_jitter.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/final_violin_plot.png +0 -0
- data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/violin_with_jitter.png +0 -0
- data/lib/galaaz/cli.rb +47 -3
- data/logos/icon-font/README.md +27 -0
- data/logos/icon-font/build_font.py +130 -0
- data/logos/icon-font/galaaz-mark.svg +34 -0
- data/script/omarchy/README.md +8 -1
- data/script/omarchy/fonts/galaaz.ttf +0 -0
- data/script/omarchy/install-galaaz.sh +8 -1
- data/script/omarchy/omarchy-menu.jsonc +21 -9
- data/sty/galaaz-header.png +0 -0
- data/sty/galaaz-headers-from-p3.tex +4 -0
- data/sty/galaaz.sty +54 -23
- data/version.rb +1 -1
- metadata +30 -9
- data/blogs/galaaz_ggplot/galaaz_ggplot.log +0 -745
- data/blogs/gknit/gknit_files/gknit_files/figure-latex/bubble-1.png +0 -0
- data/blogs/manual/manual.log +0 -1530
- data/blogs/manual/manual_files/manual_files/figure-latex/bubble-1.png +0 -0
- data/blogs/nse_dplyr/nse_dplyr.log +0 -824
- data/blogs/oh_my/oh_my.log +0 -974
- data/blogs/ruby_plot/ruby_plot.log +0 -887
|
@@ -1,134 +1,151 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: "Non Standard Evaluation in dplyr with Galaaz"
|
|
3
|
+
author:
|
|
4
|
+
- "Rodrigo Botafogo"
|
|
5
|
+
- "Daniel Mossé - University of Pittsburgh"
|
|
6
|
+
tags: [Tech, Data Science, Ruby, R, JRuby, "GNU R", Galaaz, dplyr]
|
|
7
|
+
date: "10/05/2019 (narrative updated for Galaaz 2.0, 2026)"
|
|
8
|
+
output:
|
|
9
|
+
html_document:
|
|
10
|
+
self_contained: true
|
|
11
|
+
keep_md: true
|
|
12
|
+
toc: true
|
|
13
|
+
toc_float: true
|
|
14
|
+
toc_depth: 2
|
|
15
|
+
number_sections: true
|
|
16
|
+
includes:
|
|
17
|
+
before_body: _logo_before_body.html
|
|
18
|
+
pdf_document:
|
|
19
|
+
includes:
|
|
20
|
+
in_header:
|
|
21
|
+
- "../../sty/galaaz.sty"
|
|
22
|
+
- "../../sty/galaaz-headers-from-p3.tex"
|
|
23
|
+
keep_tex: yes
|
|
24
|
+
number_sections: yes
|
|
25
|
+
toc: true
|
|
26
|
+
toc_depth: 2
|
|
27
|
+
md_document:
|
|
28
|
+
variant: markdown_github
|
|
29
|
+
fontsize: 11pt
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
|
|
33
|
+
|
|
34
|
+
|
|
35
|
+
|
|
1
36
|
# Introduction
|
|
2
37
|
|
|
3
|
-
According to Steven Sagaert’s answer on Quora about “Is programming
|
|
4
|
-
language R overrated?”:
|
|
38
|
+
According to Steven Sagaert’s answer on Quora about “Is programming language R overrated?”:
|
|
5
39
|
|
|
6
|
-
> R is a sophisticated language with an unusual (i.e.
|
|
7
|
-
>
|
|
8
|
-
>
|
|
40
|
+
> R is a sophisticated language with an unusual (i.e. non-mainstream) set of features. It‘s
|
|
41
|
+
> an impure functional programming language with sophisticated metaprogramming and 3
|
|
42
|
+
> different OO systems.
|
|
9
43
|
|
|
10
|
-
> Just like common lisp you can completely customise how things work via
|
|
11
|
-
>
|
|
12
|
-
>
|
|
13
|
-
> syntax for dplyr.
|
|
44
|
+
> Just like common lisp you can completely customise how things work via metaprogramming.
|
|
45
|
+
> The biggest example is the tidyverse: by creating it’s own evaluation system (tidyeval)
|
|
46
|
+
> was able to create a custom syntax for dplyr.
|
|
14
47
|
|
|
15
|
-
> Mastering R (the language) and its ecosystem is not a matter of weeks
|
|
16
|
-
>
|
|
48
|
+
> Mastering R (the language) and its ecosystem is not a matter of weeks or months but
|
|
49
|
+
> takes years. The rabbit hole goes pretty deep…
|
|
17
50
|
|
|
18
|
-
Although a highly configurable language can give programmers a great
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
deadline, not necessarily for building large applications.
|
|
51
|
+
Although a highly configurable language can give programmers a great deal of power,
|
|
52
|
+
it can also take years to master—as noted above. Programming with _dplyr_, for instance,
|
|
53
|
+
means learning evaluation rules that are not always approachable for **statisticians and
|
|
54
|
+
analysts who are not full-time software engineers**. That is not a criticism: R was **built**
|
|
55
|
+
for **statisticians** who need trustworthy results on a deadline, not necessarily for building
|
|
56
|
+
large applications.
|
|
25
57
|
|
|
26
|
-
**Unfortunately**, when such a user moves on to more **sophisticated**
|
|
27
|
-
|
|
58
|
+
**Unfortunately**, when such a user moves on to more **sophisticated** programming patterns,
|
|
59
|
+
the learning curve can become a real hurdle.
|
|
28
60
|
|
|
29
|
-
In this post we will see how to program with
|
|
30
|
-
|
|
61
|
+
In this post we will see how to program with _dplyr_ in Galaaz and how Ruby can simplify
|
|
62
|
+
the learning curve of mastering _dplyr_ coding.
|
|
31
63
|
|
|
32
64
|
# But first, what is Galaaz??
|
|
33
65
|
|
|
34
|
-
Galaaz is a system for tightly coupling Ruby and R.
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
libraries for data science, statistics, scientific plotting and machine
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
process so that Ruby can call **dplyr** and the rest of the tidyverse as
|
|
64
|
-
if they were part of the same workflow. An **earlier** Galaaz line of
|
|
65
|
-
work used Oracle’s **GraalVM** with **TruffleRuby** and **FastR** in a
|
|
66
|
-
single runtime; that approach is **no longer** the supported stack—see
|
|
67
|
-
the project **manual** for setup, **`bin/galaaz-jruby`**, and **gKnit**.
|
|
66
|
+
Galaaz is a system for tightly coupling Ruby and R. Ruby is a powerful language, with
|
|
67
|
+
a large community, a very large set of libraries and great for web development. It is also
|
|
68
|
+
easy to learn. However,
|
|
69
|
+
it lacks libraries for data science, statistics, scientific plotting and machine learning.
|
|
70
|
+
On the other hand, R is considered one of the most powerful languages for solving all of the
|
|
71
|
+
above problems. **Python** is a strong competitor, with NumPy, pandas, SciPy, scikit-learn,
|
|
72
|
+
and **many thousands** of other packages on PyPI. We will not dwell on R **versus** Python here:
|
|
73
|
+
both are excellent languages with different strengths.
|
|
74
|
+
Our interest is to bring to yet another excellent language, Ruby, the data science libraries
|
|
75
|
+
that it lacks.
|
|
76
|
+
|
|
77
|
+
With Galaaz we do not intend to re-implement any of the scientific libraries in R. However, we
|
|
78
|
+
allow for very tight coupling between the two languages to the point that the Ruby
|
|
79
|
+
developer does not need to know that there is an R engine running. Also, from the point of
|
|
80
|
+
view of the R user/developer, Galaaz looks a lot like R, with just minor syntactic difference,
|
|
81
|
+
so there is almost no learning curve for the R developer. And as we will see in this
|
|
82
|
+
post that programming with _dplyr_ is easier in Galaaz than in R.
|
|
83
|
+
|
|
84
|
+
R users are probably quite knowledgeable about _dplyr_. For the Ruby developer, _dplyr_ and
|
|
85
|
+
the _tidyverse_ libraries are a set of libraries for data manipulation in R, developed by
|
|
86
|
+
Hadley Wickham, Chief Scientist at Posit (formerly RStudio) and a prolific R coder and writer.
|
|
87
|
+
|
|
88
|
+
For the coupling of Ruby and R, **Galaaz 2.0** uses **[JRuby](https://www.jruby.org/)** (Ruby on the JVM)
|
|
89
|
+
together with **GNU R**. A **bridge** sends expressions and data between Ruby and an R process so that
|
|
90
|
+
Ruby can call **dplyr** and the rest of the tidyverse as if they were part of the same workflow.
|
|
91
|
+
An **earlier** Galaaz line of work used Oracle’s **GraalVM** with **TruffleRuby** and **FastR** in a single
|
|
92
|
+
runtime; that approach is **no longer** the supported stack—see the project **manual** for setup,
|
|
93
|
+
**`bin/galaaz-jruby`**, and **gKnit**.
|
|
94
|
+
|
|
68
95
|
|
|
69
96
|
# Tidyverse and dplyr
|
|
70
97
|
|
|
71
|
-
In [What is the
|
|
72
|
-
tidyverse
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
>
|
|
76
|
-
>
|
|
77
|
-
>
|
|
78
|
-
>
|
|
79
|
-
>
|
|
80
|
-
>
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
>
|
|
89
|
-
>
|
|
90
|
-
|
|
91
|
-
>
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
> 5. arrange() changes the ordering of the rows.
|
|
97
|
-
|
|
98
|
-
Very often R is used interactively and users use *dplyr* to manipulate a
|
|
99
|
-
single dataset without programming. When users want to replicate their
|
|
100
|
-
work for multiple datasets, programming becomes necessary.
|
|
98
|
+
In [What is the tidyverse?](https://rviews.rstudio.com/2017/06/08/what-is-the-tidyverse/) the
|
|
99
|
+
tidyverse is explained as follows:
|
|
100
|
+
|
|
101
|
+
> The tidyverse is a coherent system of packages for data manipulation, exploration and
|
|
102
|
+
> visualization that share a common design philosophy. These were mostly developed by
|
|
103
|
+
> Hadley Wickham himself, but they are now being expanded by several contributors. Tidyverse
|
|
104
|
+
> packages are intended to make statisticians and data scientists more productive by
|
|
105
|
+
> guiding them through workflows that facilitate communication, and result in reproducible
|
|
106
|
+
> work products. Fundamentally, the tidyverse is about the connections between the tools
|
|
107
|
+
> that make the workflow possible.
|
|
108
|
+
|
|
109
|
+
_dplyr_ is one of the many packages that are part of the tidyverse. It is:
|
|
110
|
+
|
|
111
|
+
> a grammar of data manipulation, providing a consistent set of verbs that help you solve
|
|
112
|
+
> the most common data manipulation challenges:
|
|
113
|
+
|
|
114
|
+
> 1. mutate() adds new variables that are functions of existing variables
|
|
115
|
+
> 2. select() picks variables based on their names.
|
|
116
|
+
> 3. filter() picks cases based on their values.
|
|
117
|
+
> 4. summarise() reduces multiple values down to a single summary.
|
|
118
|
+
> 5. arrange() changes the ordering of the rows.
|
|
119
|
+
|
|
120
|
+
Very often R is used interactively and users use _dplyr_ to manipulate a single dataset
|
|
121
|
+
without programming. When users want to replicate their work for
|
|
122
|
+
multiple datasets, programming becomes necessary.
|
|
101
123
|
|
|
102
124
|
# Programming with dplyr
|
|
103
125
|
|
|
104
|
-
In the vignette [
|
|
105
|
-
|
|
106
|
-
Wickham states:
|
|
126
|
+
In the vignette ["Programming with dplyr"](https://dplyr.tidyverse.org/articles/programming.html),
|
|
127
|
+
Hadley Wickham states:
|
|
107
128
|
|
|
108
|
-
> Most dplyr functions use non-standard evaluation (NSE). This is a
|
|
109
|
-
>
|
|
110
|
-
>
|
|
111
|
-
>
|
|
112
|
-
> code:
|
|
129
|
+
> Most dplyr functions use non-standard evaluation (NSE). This is a catch-all term that
|
|
130
|
+
> means they don’t follow the usual R rules of evaluation. Instead, they capture the
|
|
131
|
+
> expression that you typed and evaluate it in a custom way. This has two main
|
|
132
|
+
> benefits for dplyr code:
|
|
113
133
|
|
|
114
|
-
> Operations on data frames can be expressed succinctly because you
|
|
115
|
-
>
|
|
116
|
-
>
|
|
117
|
-
> df$y ==2 & df$z == 3, \].
|
|
134
|
+
> Operations on data frames can be expressed succinctly because you don’t need to repeat
|
|
135
|
+
> the name of the data frame. For example, you can write filter(df, x == 1, y == 2, z == 3)
|
|
136
|
+
> instead of df[df\$x == 1 & df\$y ==2 & df\$z == 3, ].
|
|
118
137
|
|
|
119
|
-
> dplyr can choose to compute results in a different way to base R. This
|
|
120
|
-
>
|
|
121
|
-
>
|
|
122
|
-
> do.
|
|
138
|
+
> dplyr can choose to compute results in a different way to base R. This is important for
|
|
139
|
+
> database backends because dplyr itself doesn’t do any work, but instead generates the SQL
|
|
140
|
+
> that tells the database what to do.
|
|
123
141
|
|
|
124
142
|
But then he goes on:
|
|
125
143
|
|
|
126
|
-
> Unfortunately these benefits do not come for free. There are two main
|
|
127
|
-
|
|
144
|
+
> Unfortunately these benefits do not come for free. There are two main drawbacks:
|
|
145
|
+
|
|
146
|
+
> Most dplyr arguments are not referentially transparent. That means you can’t replace a value
|
|
147
|
+
> with a seemingly equivalent object that you’ve defined elsewhere. In other words, this code:
|
|
128
148
|
|
|
129
|
-
> Most dplyr arguments are not referentially transparent. That means you
|
|
130
|
-
> can’t replace a value with a seemingly equivalent object that you’ve
|
|
131
|
-
> defined elsewhere. In other words, this code:
|
|
132
149
|
|
|
133
150
|
``` r
|
|
134
151
|
df <- data.frame(x = 1:3, y = 3:1)
|
|
@@ -138,78 +155,65 @@ print(filter(df, x == 1))
|
|
|
138
155
|
#> <int> <int>
|
|
139
156
|
#> 1 1 3
|
|
140
157
|
```
|
|
141
|
-
|
|
142
158
|
> Is not equivalent to this code:
|
|
143
159
|
|
|
160
|
+
|
|
144
161
|
``` r
|
|
145
162
|
my_var <- x
|
|
146
163
|
#> Error in eval(expr, envir, enclos): object 'x' not found
|
|
147
164
|
filter(df, my_var == 1)
|
|
148
165
|
#> Error: object 'my_var' not found
|
|
149
166
|
```
|
|
167
|
+
> This makes it hard to create functions with arguments that change how dplyr verbs are computed.
|
|
150
168
|
|
|
151
|
-
|
|
152
|
-
|
|
169
|
+
As a result of this, programming with _dplyr_ requires learning a set of new ideas and concepts.
|
|
170
|
+
In this vignette Hadley goes on showing how to program ever more difficult problems with _dplyr_,
|
|
171
|
+
showing the problems it faces and the new concepts needed to solve them.
|
|
153
172
|
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
program ever more difficult problems with *dplyr*, showing the problems
|
|
157
|
-
it faces and the new concepts needed to solve them.
|
|
173
|
+
In this blog, we will look at all the problems presented by Harley on the vignette and show how
|
|
174
|
+
those same problems can be solved using Galaaz and the Ruby language.
|
|
158
175
|
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
176
|
+
This blog is organized as follows: first we show how to write expressions using Galaaz.
|
|
177
|
+
Expressions are a fundamental concept in _dplyr_ and are not part of basic Ruby. We extend
|
|
178
|
+
the Ruby language create a manipulate expressions that will be used by _dplyr_ functions.
|
|
162
179
|
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
basic Ruby. We extend the Ruby language create a manipulate expressions
|
|
167
|
-
that will be used by *dplyr* functions.
|
|
180
|
+
Then we show very succintly how Ruby and R can be integrated and how R functions are
|
|
181
|
+
transparently called from Ruby. Galaaz [user manual](https://github.com/rbotafogo/galaaz/wiki)
|
|
182
|
+
(still in development) goes in much deeper detail about this integration.
|
|
168
183
|
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
goes in much deeper detail about this integration.
|
|
184
|
+
Next in section "Data manipulation wiht _dplyr_" we go through all the problems on the
|
|
185
|
+
_dplyr_ vignette and look at how they are solved in Galaaz. We then discuss why programming
|
|
186
|
+
with Galaaz and _dplyr_ is easier than programming with _dplyr_ in plain R.
|
|
173
187
|
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
Galaaz. We then discuss why programming with Galaaz and *dplyr* is
|
|
177
|
-
easier than programming with *dplyr* in plain R.
|
|
178
|
-
|
|
179
|
-
The following section looks at another more advanced problem and shows
|
|
180
|
-
that Galaaz can still handle it without any difficulty. We then provide
|
|
181
|
-
further reading and concluding remarks.
|
|
188
|
+
The following section looks at another more advanced problem and shows that Galaaz can still
|
|
189
|
+
handle it without any difficulty. We then provide further reading and concluding remarks.
|
|
182
190
|
|
|
183
191
|
# Writing Expressions in Galaaz
|
|
184
192
|
|
|
185
|
-
Galaaz extends Ruby to work with expressions, similar to R
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
but cannot be computed unless the value of *x* is bound to some value.
|
|
193
|
+
Galaaz extends Ruby to work with expressions, similar to R's expressions build with 'quote'
|
|
194
|
+
(base R) or 'quo' (tidyverse). Expressions in this context are like mathematical expressions or
|
|
195
|
+
formulae. For instance, in mathematics, the expression $y = sin(x)$ describes a function but cannot
|
|
196
|
+
be computed unless the value of $x$ is bound to some value.
|
|
190
197
|
|
|
191
|
-
Expressions are fundamental in
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
*dplyr* function with the expression ‘y = x \* 2’.
|
|
198
|
+
Expressions are fundamental in _dplyr_ programming as they are the input to _dplyr_ functions,
|
|
199
|
+
for instance, as we will see shortly, if a data frame has a column named 'x' and we want
|
|
200
|
+
to add another column, y, to this dataframe that has the values of 'x' times 2, then we would
|
|
201
|
+
call a _dplyr_ function with the expression 'y = x * 2'.
|
|
196
202
|
|
|
197
203
|
## A note on notation
|
|
198
204
|
|
|
199
|
-
This blog was written in Rmarkdown and automatically converted to HTML
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
blocks
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
text (PDF). Every output line from the code execution is preceded by
|
|
206
|
-
‘\##’.
|
|
205
|
+
This blog was written in Rmarkdown and automatically converted to HTML or PDF (depending on
|
|
206
|
+
where you are reading this blog) with gKnit (a tool provided by Galaaz). In Rmarkdown, it is
|
|
207
|
+
possible to write text and code blocks that are executed to generate the final report. Code
|
|
208
|
+
blocks appear inside a 'box' and the result of their execution appear either in another type
|
|
209
|
+
of 'box' with a different background (HTML) or as normal text (PDF). Every output line from
|
|
210
|
+
the code execution is preceded by '##'.
|
|
207
211
|
|
|
208
212
|
## Expressions from operators
|
|
209
213
|
|
|
210
|
-
The code below creates an expression summing two symbols. Note that :a
|
|
211
|
-
|
|
212
|
-
|
|
214
|
+
The code below creates an expression summing two symbols. Note that :a and :b are Ruby symbols and
|
|
215
|
+
are not bound to any values at the time of expression definition:
|
|
216
|
+
|
|
213
217
|
|
|
214
218
|
``` ruby
|
|
215
219
|
begin
|
|
@@ -221,36 +225,41 @@ rescue => e
|
|
|
221
225
|
end
|
|
222
226
|
```
|
|
223
227
|
|
|
224
|
-
|
|
225
|
-
|
|
228
|
+
```
|
|
229
|
+
## NoMethodError: undefined method '+' for an instance of Symbol
|
|
230
|
+
```
|
|
226
231
|
In Galaaz, we can build any complex mathematical expression such as:
|
|
227
232
|
|
|
233
|
+
|
|
228
234
|
``` ruby
|
|
229
235
|
exp2 = (R[:a] + R[:b]) * 2.0 + R[:c] ** 2 / R[:z]
|
|
230
236
|
puts exp2
|
|
231
237
|
```
|
|
232
238
|
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
|
|
236
|
-
|
|
239
|
+
```
|
|
240
|
+
## a + b * 2.0 + c ^ 2L / z
|
|
241
|
+
```
|
|
242
|
+
Expressions are printed with the same format as the equivalent R expressions. The 'L' after
|
|
243
|
+
2 indicates that 2 is an integer.
|
|
237
244
|
|
|
238
|
-
The R developer should note that in R, if she writes the
|
|
239
|
-
R interpreter will convert it to float. In order to get an interger she
|
|
240
|
-
should write
|
|
241
|
-
|
|
245
|
+
The R developer should note that in R, if she writes the
|
|
246
|
+
number '2', the R interpreter will convert it to float. In order to get an interger she
|
|
247
|
+
should write '2L'. Galaaz follows Ruby notation and '2' is an integer, while '2.0' is a
|
|
248
|
+
float.
|
|
242
249
|
|
|
243
250
|
It is also possible to use inequality operators in building expressions:
|
|
244
251
|
|
|
252
|
+
|
|
245
253
|
``` ruby
|
|
246
254
|
exp3 = (R[:a] + R[:b]) >= R[:z]
|
|
247
255
|
puts exp3
|
|
248
256
|
```
|
|
249
257
|
|
|
250
|
-
|
|
258
|
+
```
|
|
259
|
+
## a + b >= z
|
|
260
|
+
```
|
|
261
|
+
Expressions' definition can also make use of normal Ruby variables without any problem:
|
|
251
262
|
|
|
252
|
-
Expressions’ definition can also make use of normal Ruby variables
|
|
253
|
-
without any problem:
|
|
254
263
|
|
|
255
264
|
``` ruby
|
|
256
265
|
x = 20
|
|
@@ -259,116 +268,133 @@ exp_var = (R[:a] + R[:b]) * x <= R[:z] - y
|
|
|
259
268
|
puts exp_var
|
|
260
269
|
```
|
|
261
270
|
|
|
262
|
-
|
|
271
|
+
```
|
|
272
|
+
## a + b * 20L <= z - 30.0
|
|
273
|
+
```
|
|
274
|
+
|
|
275
|
+
Galaaz provides both symbolic representations for operators, such as (>, <, !=) as functional
|
|
276
|
+
notation for those operators such as (.gt, .ge, etc.). So the same expression written
|
|
277
|
+
above can also be written as
|
|
263
278
|
|
|
264
|
-
Galaaz provides both symbolic representations for operators, such as
|
|
265
|
-
(\>, \<, !=) as functional notation for those operators such as (.gt,
|
|
266
|
-
.ge, etc.). So the same expression written above can also be written as
|
|
267
279
|
|
|
268
280
|
``` ruby
|
|
269
281
|
exp4 = (R[:a] + R[:b]).ge R[:z]
|
|
270
282
|
puts exp4
|
|
271
283
|
```
|
|
272
284
|
|
|
273
|
-
|
|
285
|
+
```
|
|
286
|
+
## a + b >= z
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
Two types of expressions, however, can only be created with the functional representation
|
|
290
|
+
of the operators. Those are expressions involving '==', and '='. This is the case since
|
|
291
|
+
those symbols have special meaning in Ruby and should not be redefined.
|
|
274
292
|
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
involving ‘==’, and ‘=’. This is the case since those symbols have
|
|
278
|
-
special meaning in Ruby and should not be redefined.
|
|
293
|
+
In order to write an expression involving '==' we
|
|
294
|
+
need to use the method '.eq' and for '=' we need the function '.assign':
|
|
279
295
|
|
|
280
|
-
In order to write an expression involving ‘==’ we need to use the method
|
|
281
|
-
‘.eq’ and for ‘=’ we need the function ‘.assign’:
|
|
282
296
|
|
|
283
297
|
``` ruby
|
|
284
298
|
exp5 = (R[:a] + R[:b]).eq R[:z]
|
|
285
299
|
puts exp5
|
|
286
300
|
```
|
|
287
301
|
|
|
288
|
-
|
|
302
|
+
```
|
|
303
|
+
## a + b == z
|
|
304
|
+
```
|
|
305
|
+
|
|
289
306
|
|
|
290
307
|
``` ruby
|
|
291
308
|
exp6 = R[:y].assign R[:a] + R[:b]
|
|
292
309
|
puts exp6
|
|
293
310
|
```
|
|
294
311
|
|
|
295
|
-
|
|
312
|
+
```
|
|
313
|
+
## y <- a + b
|
|
314
|
+
```
|
|
315
|
+
Users should be careful when writing expressions not to inadvertently use '==' or '=' as
|
|
316
|
+
this will generate an error, that might be a bit cryptic (in future releases of Galaza, we
|
|
317
|
+
plan to improve the error message).
|
|
296
318
|
|
|
297
|
-
Users should be careful when writing expressions not to inadvertently
|
|
298
|
-
use ‘==’ or ‘=’ as this will generate an error, that might be a bit
|
|
299
|
-
cryptic (in future releases of Galaza, we plan to improve the error
|
|
300
|
-
message).
|
|
301
319
|
|
|
302
320
|
``` ruby
|
|
303
321
|
exp_wrong = (R[:a] + R[:b]) == R[:z]
|
|
304
322
|
puts exp_wrong
|
|
305
323
|
```
|
|
306
324
|
|
|
307
|
-
|
|
308
|
-
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
325
|
+
```
|
|
326
|
+
## false
|
|
327
|
+
```
|
|
328
|
+
The problem lies with the fact that
|
|
329
|
+
when using '==' we are comparing expression (R[:a] + R[:b]) to expression R[:z] with '=='. When this
|
|
330
|
+
comparison is executed, the system tries to evaluate :a, :b and :z, and those symbols, at
|
|
331
|
+
this time, are not bound to anything giving the "object 'a' not found" message.
|
|
314
332
|
|
|
315
333
|
## Expressions with R methods
|
|
316
334
|
|
|
317
|
-
It is often necessary to create an expression that uses a method or
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
335
|
+
It is often necessary to create an expression that uses a method or function. For instance, in
|
|
336
|
+
mathematics, it's quite natural to write an expressin such as $y = sin(x)$. In this case, the
|
|
337
|
+
'sin' function is part of the expression and should not be immediately executed. When we want
|
|
338
|
+
the function to be part of the expression, we call the function preceeding it
|
|
339
|
+
by the letter E, such as 'E.sin(x)'
|
|
340
|
+
|
|
323
341
|
|
|
324
342
|
``` ruby
|
|
325
343
|
exp7 = R[:y].assign E.sin(R[:x])
|
|
326
344
|
puts exp7
|
|
327
345
|
```
|
|
328
346
|
|
|
329
|
-
|
|
347
|
+
```
|
|
348
|
+
## y <- sin(x)
|
|
349
|
+
```
|
|
350
|
+
Function expressions can also be written using '.' notation:
|
|
330
351
|
|
|
331
|
-
Function expressions can also be written using ‘.’ notation:
|
|
332
352
|
|
|
333
353
|
``` ruby
|
|
334
354
|
exp8 = R[:y].assign R[:x].sin
|
|
335
355
|
puts exp8
|
|
336
356
|
```
|
|
337
357
|
|
|
338
|
-
|
|
358
|
+
```
|
|
359
|
+
## y <- sin(x)
|
|
360
|
+
```
|
|
361
|
+
When a function has multiple arguments, the first one can be used before the '.'. For instance,
|
|
362
|
+
the R concatenate function 'c', that concatenates two or more arguments can be part of
|
|
363
|
+
an expression as:
|
|
339
364
|
|
|
340
|
-
When a function has multiple arguments, the first one can be used before
|
|
341
|
-
the ‘.’. For instance, the R concatenate function ‘c’, that concatenates
|
|
342
|
-
two or more arguments can be part of an expression as:
|
|
343
365
|
|
|
344
366
|
``` ruby
|
|
345
367
|
exp9 = R[:x].c(R[:y])
|
|
346
368
|
puts exp9
|
|
347
369
|
```
|
|
348
370
|
|
|
349
|
-
|
|
350
|
-
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
operator
|
|
371
|
+
```
|
|
372
|
+
## c(x, y)
|
|
373
|
+
```
|
|
374
|
+
Note that this gives an OO feeling to the code, as if we were saying 'x' concatenates 'y'. As a
|
|
375
|
+
side note, '.' notation can be used as the R pipe operator '%>%', but is more general than the
|
|
376
|
+
pipe.
|
|
354
377
|
|
|
355
378
|
## Evaluating an Expression
|
|
356
379
|
|
|
357
|
-
Although we are mainly focusing on expressions to pass them to
|
|
358
|
-
|
|
359
|
-
a binding.
|
|
380
|
+
Although we are mainly focusing on expressions to pass them to _dplyr_ functions, expressions
|
|
381
|
+
can be evaluated by calling function 'eval' with a binding.
|
|
360
382
|
|
|
361
383
|
A binding can be provided with a list or a data frame as shown below:
|
|
362
384
|
|
|
385
|
+
|
|
363
386
|
``` ruby
|
|
364
387
|
exp = (R[:a] + R[:b]) * 2.0 + R[:c] ** 2 / R[:z]
|
|
365
388
|
puts exp.eval(R.list(a: 10, b: 20, c: 30, z: 40))
|
|
366
389
|
```
|
|
367
390
|
|
|
368
|
-
|
|
391
|
+
```
|
|
392
|
+
## [1] 72.5
|
|
393
|
+
```
|
|
369
394
|
|
|
370
395
|
with a data frame:
|
|
371
396
|
|
|
397
|
+
|
|
372
398
|
``` ruby
|
|
373
399
|
df = R.data__frame(
|
|
374
400
|
a: R.c(1, 2, 3),
|
|
@@ -379,83 +405,90 @@ df = R.data__frame(
|
|
|
379
405
|
puts exp.eval(df)
|
|
380
406
|
```
|
|
381
407
|
|
|
382
|
-
|
|
408
|
+
```
|
|
409
|
+
## [1] 31 62 93
|
|
410
|
+
```
|
|
383
411
|
|
|
384
412
|
# Using Galaaz to call R functions
|
|
385
413
|
|
|
386
|
-
Galaaz tries to emulate as closely as possible the way R functions are
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
414
|
+
Galaaz tries to emulate as closely as possible the way R functions are called and migrating from
|
|
415
|
+
R to Galaaz should be quite easy requiring only minor syntactic changes to an R script. In
|
|
416
|
+
this post, we do not have enough space to write a complete manual on Galaaz
|
|
417
|
+
(a short manual can be found at: https://www.rubydoc.info/gems/galaaz/0.4.9), so we will
|
|
418
|
+
present only a few examples scripts using Galaaz.
|
|
419
|
+
|
|
420
|
+
Basically, to call an R function from Ruby with Galaaz, one only needs to preced the function
|
|
421
|
+
with 'R.'. For instance, to create a vector in R, the 'c' function is used. In Galaaz, a
|
|
422
|
+
vector can be created by using 'R.c':
|
|
392
423
|
|
|
393
|
-
Basically, to call an R function from Ruby with Galaaz, one only needs
|
|
394
|
-
to preced the function with ‘R.’. For instance, to create a vector in R,
|
|
395
|
-
the ‘c’ function is used. In Galaaz, a vector can be created by using
|
|
396
|
-
‘R.c’:
|
|
397
424
|
|
|
398
425
|
``` ruby
|
|
399
426
|
vec = R.c(1.0, 2, 3)
|
|
400
427
|
puts vec
|
|
401
428
|
```
|
|
402
429
|
|
|
403
|
-
|
|
430
|
+
```
|
|
431
|
+
## [1] 1 2 3
|
|
432
|
+
```
|
|
433
|
+
A list is created in R with the 'list' function, so in Galaaz we do:
|
|
404
434
|
|
|
405
|
-
A list is created in R with the ‘list’ function, so in Galaaz we do:
|
|
406
435
|
|
|
407
436
|
``` ruby
|
|
408
437
|
list = R.list(a: 1.0, b: 2, c: 3)
|
|
409
438
|
puts list
|
|
410
439
|
```
|
|
411
440
|
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
|
|
416
|
-
|
|
417
|
-
|
|
418
|
-
|
|
419
|
-
|
|
441
|
+
```
|
|
442
|
+
## $a
|
|
443
|
+
## [1] 1
|
|
444
|
+
##
|
|
445
|
+
## $b
|
|
446
|
+
## [1] 2
|
|
447
|
+
##
|
|
448
|
+
## $c
|
|
449
|
+
## [1] 3
|
|
450
|
+
```
|
|
451
|
+
Note that we can use named arguments in our list. The same code in R would be:
|
|
420
452
|
|
|
421
|
-
Note that we can use named arguments in our list. The same code in R
|
|
422
|
-
would be:
|
|
423
453
|
|
|
424
454
|
``` r
|
|
425
455
|
lst = list(a = 1, b = 2L, c = 3L)
|
|
426
456
|
print(lst)
|
|
427
457
|
```
|
|
428
458
|
|
|
429
|
-
|
|
430
|
-
|
|
431
|
-
|
|
432
|
-
|
|
433
|
-
|
|
434
|
-
|
|
435
|
-
|
|
436
|
-
|
|
459
|
+
```
|
|
460
|
+
## $a
|
|
461
|
+
## [1] 1
|
|
462
|
+
##
|
|
463
|
+
## $b
|
|
464
|
+
## [1] 2
|
|
465
|
+
##
|
|
466
|
+
## $c
|
|
467
|
+
## [1] 3
|
|
468
|
+
```
|
|
469
|
+
Now, let's say that 'x' is an angle of 45$^\circ$ and we acttually want to create
|
|
470
|
+
the expression $y = sin(45^\circ)$, which is $y = 0.850...$. In this case,
|
|
471
|
+
we will use 'R.sin':
|
|
437
472
|
|
|
438
|
-
Now, let’s say that ‘x’ is an angle of 45<sup>∘</sup> and we acttually
|
|
439
|
-
want to create the expression *y* = *s**i**n*(45<sup>∘</sup>), which is
|
|
440
|
-
*y* = 0.850.... In this case, we will use ‘R.sin’:
|
|
441
473
|
|
|
442
474
|
``` ruby
|
|
443
475
|
exp10 = R[:y].assign R.sin(45)
|
|
444
476
|
puts exp10
|
|
445
477
|
```
|
|
446
478
|
|
|
447
|
-
|
|
479
|
+
```
|
|
480
|
+
## y <- 0.850903524534118
|
|
481
|
+
```
|
|
482
|
+
|
|
483
|
+
# Data manipulation wiht _dplyr_
|
|
448
484
|
|
|
449
|
-
|
|
485
|
+
In this section we will give a brief tour _dplyr_'s usage in Galaaz and how to manipulate
|
|
486
|
+
data in Ruby with it. This section will follow [_dplyr_'s vignette](https://dplyr.tidyverse.org/articles/dplyr.html) that explores the nycflights13 data set. This dataset contains all 336776
|
|
487
|
+
flights that departed from New York City in 2013. The data comes from the US Bureau of
|
|
488
|
+
Transportation Statistics.
|
|
450
489
|
|
|
451
|
-
|
|
452
|
-
how to manipulate data in Ruby with it. This section will follow
|
|
453
|
-
[*dplyr*’s vignette](https://dplyr.tidyverse.org/articles/dplyr.html)
|
|
454
|
-
that explores the nycflights13 data set. This dataset contains all
|
|
455
|
-
336776 flights that departed from New York City in 2013. The data comes
|
|
456
|
-
from the US Bureau of Transportation Statistics.
|
|
490
|
+
Let's start by taking a look at this dataset:
|
|
457
491
|
|
|
458
|
-
Let’s start by taking a look at this dataset:
|
|
459
492
|
|
|
460
493
|
``` ruby
|
|
461
494
|
R.library('nycflights13')
|
|
@@ -465,137 +498,141 @@ puts ~R[:flights].dim
|
|
|
465
498
|
~R[:flights].str
|
|
466
499
|
```
|
|
467
500
|
|
|
468
|
-
|
|
469
|
-
|
|
501
|
+
```
|
|
502
|
+
## ~(dim(flights))
|
|
503
|
+
## <environment: 0x5cd09170df98>
|
|
504
|
+
```
|
|
470
505
|
|
|
471
|
-
Now, let
|
|
472
|
-
|
|
473
|
-
filter
|
|
474
|
-
function
|
|
475
|
-
|
|
476
|
-
|
|
506
|
+
Now, let's use a first verb of _dplyr_: 'filter'. This verb, obviously, will filter the data
|
|
507
|
+
by the given expression. In the next block, we filter by columns 'month' and 'day'. The
|
|
508
|
+
first argument to the filter function is symbol ':flights'. A Ruby symbol, when given to
|
|
509
|
+
an R function will convert to the R variable of the same name, in this case 'flights', that
|
|
510
|
+
holds the nycflights13 data frame.
|
|
511
|
+
|
|
512
|
+
The second and third arguments are expressions that will be used by the filter function to
|
|
513
|
+
filter by columns, looking for entries in which the month and day are equal to 1.
|
|
477
514
|
|
|
478
|
-
The second and third arguments are expressions that will be used by the
|
|
479
|
-
filter function to filter by columns, looking for entries in which the
|
|
480
|
-
month and day are equal to 1.
|
|
481
515
|
|
|
482
516
|
``` ruby
|
|
483
517
|
puts R.filter(:flights, (R[:month].eq 1), (R[:day].eq 1))
|
|
484
518
|
```
|
|
485
519
|
|
|
486
|
-
|
|
487
|
-
|
|
488
|
-
|
|
489
|
-
|
|
490
|
-
|
|
491
|
-
|
|
492
|
-
|
|
493
|
-
|
|
494
|
-
|
|
495
|
-
|
|
496
|
-
|
|
497
|
-
|
|
498
|
-
|
|
499
|
-
|
|
500
|
-
|
|
501
|
-
|
|
502
|
-
|
|
503
|
-
|
|
504
|
-
|
|
505
|
-
|
|
506
|
-
|
|
507
|
-
|
|
508
|
-
|
|
509
|
-
|
|
510
|
-
|
|
511
|
-
|
|
520
|
+
```
|
|
521
|
+
## # A tibble: 842 × 19
|
|
522
|
+
## year month day dep_time sched_dep_time dep_delay arr_time
|
|
523
|
+
## <int> <int> <int> <int> <int> <dbl> <int>
|
|
524
|
+
## 1 2013 1 1 517 515 2 830
|
|
525
|
+
## 2 2013 1 1 533 529 4 850
|
|
526
|
+
## 3 2013 1 1 542 540 2 923
|
|
527
|
+
## 4 2013 1 1 544 545 -1 1004
|
|
528
|
+
## 5 2013 1 1 554 600 -6 812
|
|
529
|
+
## 6 2013 1 1 554 558 -4 740
|
|
530
|
+
## 7 2013 1 1 555 600 -5 913
|
|
531
|
+
## 8 2013 1 1 557 600 -3 709
|
|
532
|
+
## 9 2013 1 1 557 600 -3 838
|
|
533
|
+
## 10 2013 1 1 558 600 -2 753
|
|
534
|
+
## # ℹ 832 more rows
|
|
535
|
+
## # ℹ 12 more variables: sched_arr_time <int>, arr_delay <dbl>,
|
|
536
|
+
## # carrier <chr>, flight <int>, tailnum <chr>, origin <chr>,
|
|
537
|
+
## # dest <chr>, air_time <dbl>, distance <dbl>, hour <dbl>,
|
|
538
|
+
## # minute <dbl>, time_hour <dttm>
|
|
539
|
+
```
|
|
540
|
+
|
|
541
|
+
|
|
542
|
+
## Programming with _dplyr_: problems and how to solve them in Galaaz
|
|
543
|
+
|
|
544
|
+
In this section we look at the list of problems that Hadley describes in the "Programming with dplyr"
|
|
545
|
+
vignette and show how those problems are solved and coded with Galaaz. Readers interested in
|
|
546
|
+
how those problems are treated in _dplyr_ should read the vignette and use it as a comparison with
|
|
547
|
+
this blog.
|
|
512
548
|
|
|
513
549
|
## Filtering using expressions
|
|
514
550
|
|
|
515
|
-
Now that we know how to write expressions and call R functions, let
|
|
516
|
-
|
|
517
|
-
frame.
|
|
518
|
-
|
|
519
|
-
|
|
520
|
-
|
|
521
|
-
‘R.data\_\_frame’:
|
|
551
|
+
Now that we know how to write expressions and call R functions, let's do some data manipulation in
|
|
552
|
+
Galaaz. Let's first start by creating a data frame. In R, the 'data.frame' function creates a
|
|
553
|
+
data frame. In Ruby, writing 'data.frame' will not parse as a single object. To call R
|
|
554
|
+
functions that have a '.' in them, we need to substitute the '.' with '__'. So, method
|
|
555
|
+
'data.frame' in R, is called in Galaaz as 'R.data\_\_frame':
|
|
556
|
+
|
|
522
557
|
|
|
523
558
|
``` ruby
|
|
524
559
|
df = R.data__frame(x: (1..3), y: (3..1))
|
|
525
560
|
puts df
|
|
526
561
|
```
|
|
527
562
|
|
|
528
|
-
|
|
529
|
-
|
|
530
|
-
|
|
531
|
-
|
|
563
|
+
```
|
|
564
|
+
## x y
|
|
565
|
+
## 1 1 3
|
|
566
|
+
## 2 2 2
|
|
567
|
+
## 3 3 1
|
|
568
|
+
```
|
|
532
569
|
|
|
533
|
-
|
|
534
|
-
|
|
535
|
-
|
|
570
|
+
_dplyr_ provides the 'filter' function, that filters data in a data brame. The 'filter'
|
|
571
|
+
function can be called on this data frame either by using 'R.filter(df, ...)' or
|
|
572
|
+
by using dot notation.
|
|
536
573
|
|
|
537
|
-
|
|
574
|
+
-------FIX---------
|
|
575
|
+
|
|
576
|
+
We prefer to use dot notation as shown below. The argument to 'filter' should be an
|
|
577
|
+
expression. Note that if we gave to filter a Ruby expression such as
|
|
578
|
+
'x == 1', we would get an error, since there is no variable 'x' defined and if 'x' was a variable
|
|
579
|
+
then 'x == 1' would either be 'true' or 'false'. Our goal is to filter our data frame returning
|
|
580
|
+
all rows in which the 'x' value is equal to 1. To express this we want: 'R[:x].eq 1', where :x will
|
|
581
|
+
be interpreted by filter as the 'x' column.
|
|
538
582
|
|
|
539
|
-
We prefer to use dot notation as shown below. The argument to ‘filter’
|
|
540
|
-
should be an expression. Note that if we gave to filter a Ruby
|
|
541
|
-
expression such as ‘x == 1’, we would get an error, since there is no
|
|
542
|
-
variable ‘x’ defined and if ‘x’ was a variable then ‘x == 1’ would
|
|
543
|
-
either be ‘true’ or ‘false’. Our goal is to filter our data frame
|
|
544
|
-
returning all rows in which the ‘x’ value is equal to 1. To express this
|
|
545
|
-
we want: ‘R\[:x\].eq 1’, where :x will be interpreted by filter as the
|
|
546
|
-
‘x’ column.
|
|
547
583
|
|
|
548
584
|
``` ruby
|
|
549
585
|
puts df.filter(R[:x].eq 1)
|
|
550
586
|
```
|
|
551
587
|
|
|
552
|
-
|
|
553
|
-
|
|
588
|
+
```
|
|
589
|
+
## x y
|
|
590
|
+
## 1 1 3
|
|
591
|
+
```
|
|
592
|
+
In R, and when coding with 'tidyverse', arguments to a function are usually not
|
|
593
|
+
*referencially transparent*. That is, you can’t replace a value with a seemingly equivalent
|
|
594
|
+
object that you’ve defined elsewhere. In other words, this code
|
|
554
595
|
|
|
555
|
-
In R, and when coding with ‘tidyverse’, arguments to a function are
|
|
556
|
-
usually not *referencially transparent*. That is, you can’t replace a
|
|
557
|
-
value with a seemingly equivalent object that you’ve defined elsewhere.
|
|
558
|
-
In other words, this code
|
|
559
596
|
|
|
560
597
|
``` r
|
|
561
598
|
my_var <- x
|
|
562
599
|
filter(df, my_var == 1)
|
|
563
600
|
```
|
|
601
|
+
Generates the following error: "object 'x' not found.
|
|
564
602
|
|
|
565
|
-
|
|
603
|
+
However, in Galaaz, arguments are referencially transparent as can be seen by the
|
|
604
|
+
code below. Note initially that 'my_var = R[:x]' will not give the error "object 'x' not found"
|
|
605
|
+
since ':x' is treated as an expression and assigned to my\_var. Then when doing (my\_var.eq 1),
|
|
606
|
+
my\_var is a variable that resolves to ':x' and it becomes equivalent to (R[:x].eq 1) which is
|
|
607
|
+
what we want.
|
|
566
608
|
|
|
567
|
-
However, in Galaaz, arguments are referencially transparent as can be
|
|
568
|
-
seen by the code below. Note initially that ‘my_var = R\[:x\]’ will not
|
|
569
|
-
give the error “object ‘x’ not found” since ‘:x’ is treated as an
|
|
570
|
-
expression and assigned to my_var. Then when doing (my_var.eq 1), my_var
|
|
571
|
-
is a variable that resolves to ‘:x’ and it becomes equivalent to
|
|
572
|
-
(R\[:x\].eq 1) which is what we want.
|
|
573
609
|
|
|
574
610
|
``` ruby
|
|
575
611
|
my_var = R[:x]
|
|
576
612
|
puts df.filter(my_var.eq 1)
|
|
577
613
|
```
|
|
578
614
|
|
|
579
|
-
|
|
580
|
-
|
|
581
|
-
|
|
615
|
+
```
|
|
616
|
+
## x y
|
|
617
|
+
## 1 1 3
|
|
618
|
+
```
|
|
582
619
|
As stated by Hadley
|
|
583
620
|
|
|
584
|
-
> dplyr code is ambiguous. Depending on what variables are defined
|
|
585
|
-
>
|
|
621
|
+
> dplyr code is ambiguous. Depending on what variables are defined where,
|
|
622
|
+
> filter(df, x == y) could be equivalent to any of:
|
|
586
623
|
|
|
587
|
-
|
|
588
|
-
|
|
589
|
-
|
|
590
|
-
|
|
624
|
+
```
|
|
625
|
+
df[df$x == df$y, ]
|
|
626
|
+
df[df$x == y, ]
|
|
627
|
+
df[x == df$y, ]
|
|
628
|
+
df[x == y, ]
|
|
629
|
+
```
|
|
630
|
+
In galaaz this ambiguity does not exist, filter(df, x.eq y) is not a valid expression as
|
|
631
|
+
expressions are build with symbols. In doing filter(df, R[:x].eq y) we are looking for elements
|
|
632
|
+
of the 'x' column that are equal to a previously defined y variable. Finally in
|
|
633
|
+
filter(df, R[:x].eq R[:y]) we are looking for elements in which the 'x' column value is equal to
|
|
634
|
+
the 'y' column value. This can be seen in the following two chunks of code:
|
|
591
635
|
|
|
592
|
-
In galaaz this ambiguity does not exist, filter(df, x.eq y) is not a
|
|
593
|
-
valid expression as expressions are build with symbols. In doing
|
|
594
|
-
filter(df, R\[:x\].eq y) we are looking for elements of the ‘x’ column
|
|
595
|
-
that are equal to a previously defined y variable. Finally in filter(df,
|
|
596
|
-
R\[:x\].eq R\[:y\]) we are looking for elements in which the ‘x’ column
|
|
597
|
-
value is equal to the ‘y’ column value. This can be seen in the
|
|
598
|
-
following two chunks of code:
|
|
599
636
|
|
|
600
637
|
``` ruby
|
|
601
638
|
y = 1
|
|
@@ -605,8 +642,11 @@ x = 2
|
|
|
605
642
|
puts df.filter(R[:x].eq R[:y])
|
|
606
643
|
```
|
|
607
644
|
|
|
608
|
-
|
|
609
|
-
|
|
645
|
+
```
|
|
646
|
+
## x y
|
|
647
|
+
## 1 2 2
|
|
648
|
+
```
|
|
649
|
+
|
|
610
650
|
|
|
611
651
|
``` ruby
|
|
612
652
|
# looking for values where the 'x' column is equal to the 'y' variable
|
|
@@ -614,35 +654,37 @@ puts df.filter(R[:x].eq R[:y])
|
|
|
614
654
|
puts df.filter(R[:x].eq y)
|
|
615
655
|
```
|
|
616
656
|
|
|
617
|
-
|
|
618
|
-
|
|
619
|
-
|
|
657
|
+
```
|
|
658
|
+
## x y
|
|
659
|
+
## 1 1 3
|
|
660
|
+
```
|
|
620
661
|
## Writing a function that applies to different data sets
|
|
621
662
|
|
|
622
|
-
Let
|
|
623
|
-
|
|
624
|
-
|
|
625
|
-
column ‘a’ plus ‘x’.
|
|
626
|
-
|
|
627
|
-
Here is the intended behaviour using the ‘mutate’ function of ‘dplyr’:
|
|
663
|
+
Let's suppose that we want to write a function that receives as the first argument a data frame
|
|
664
|
+
and as second argument an expression that adds a column to the data frame that is equal to the
|
|
665
|
+
sum of elements in column 'a' plus 'x'.
|
|
628
666
|
|
|
629
|
-
|
|
630
|
-
mutate(df2, y = a + x)
|
|
631
|
-
mutate(df3, y = a + x)
|
|
632
|
-
mutate(df4, y = a + x)
|
|
667
|
+
Here is the intended behaviour using the 'mutate' function of 'dplyr':
|
|
633
668
|
|
|
669
|
+
```
|
|
670
|
+
mutate(df1, y = a + x)
|
|
671
|
+
mutate(df2, y = a + x)
|
|
672
|
+
mutate(df3, y = a + x)
|
|
673
|
+
mutate(df4, y = a + x)
|
|
674
|
+
```
|
|
634
675
|
The naive approach to writing an R function to solve this problem is:
|
|
635
676
|
|
|
636
|
-
|
|
637
|
-
|
|
638
|
-
|
|
677
|
+
```
|
|
678
|
+
mutate_y <- function(df) {
|
|
679
|
+
mutate(df, y = a + x)
|
|
680
|
+
}
|
|
681
|
+
```
|
|
682
|
+
Unfortunately, in R, this function can fail silently if one of the variables isn’t present
|
|
683
|
+
in the data frame, but is present in the global environment. We will not go through here how
|
|
684
|
+
to solve this problem in R.
|
|
639
685
|
|
|
640
|
-
|
|
641
|
-
variables isn’t present in the data frame, but is present in the global
|
|
642
|
-
environment. We will not go through here how to solve this problem in R.
|
|
686
|
+
In Galaaz the method mutate_y below will work fine and will never fail silently.
|
|
643
687
|
|
|
644
|
-
In Galaaz the method mutate_y below will work fine and will never fail
|
|
645
|
-
silently.
|
|
646
688
|
|
|
647
689
|
``` ruby
|
|
648
690
|
def mutate_y(df)
|
|
@@ -651,23 +693,25 @@ def mutate_y(df)
|
|
|
651
693
|
df.mutate(y: R[:a] + R[:x])
|
|
652
694
|
end
|
|
653
695
|
```
|
|
696
|
+
Here we create a data frame that has only one column named 'x':
|
|
654
697
|
|
|
655
|
-
Here we create a data frame that has only one column named ‘x’:
|
|
656
698
|
|
|
657
699
|
``` ruby
|
|
658
700
|
df1 = R.data__frame(x: (1..3))
|
|
659
701
|
puts df1
|
|
660
702
|
```
|
|
661
703
|
|
|
662
|
-
|
|
663
|
-
|
|
664
|
-
|
|
665
|
-
|
|
704
|
+
```
|
|
705
|
+
## x
|
|
706
|
+
## 1 1
|
|
707
|
+
## 2 2
|
|
708
|
+
## 3 3
|
|
709
|
+
```
|
|
710
|
+
|
|
711
|
+
Note that method mutate_y will fail independetly from the fact that variable 'a' is defined and
|
|
712
|
+
in the scope of the method. Variable 'a' has no relationship with the symbol `R[:a]` used in the
|
|
713
|
+
definition of 'mutate\_y' above:
|
|
666
714
|
|
|
667
|
-
Note that method mutate_y will fail independetly from the fact that
|
|
668
|
-
variable ‘a’ is defined and in the scope of the method. Variable ‘a’ has
|
|
669
|
-
no relationship with the symbol `R[:a]` used in the definition of
|
|
670
|
-
‘mutate_y’ above:
|
|
671
715
|
|
|
672
716
|
``` ruby
|
|
673
717
|
a = 10
|
|
@@ -679,17 +723,18 @@ rescue => e
|
|
|
679
723
|
end
|
|
680
724
|
```
|
|
681
725
|
|
|
682
|
-
|
|
683
|
-
|
|
684
|
-
|
|
685
|
-
|
|
726
|
+
```
|
|
727
|
+
## NewBridge::SessionClient::RProcessError: Error: ℹ In argument: `y = a + x`.
|
|
728
|
+
## Caused by error:
|
|
729
|
+
## ! object 'a' not found
|
|
730
|
+
```
|
|
686
731
|
## Different expressions
|
|
687
732
|
|
|
688
|
-
Let
|
|
689
|
-
|
|
690
|
-
|
|
691
|
-
|
|
692
|
-
|
|
733
|
+
Let's move to the next problem as presented by Hadley where trying to write a function in R
|
|
734
|
+
that will receive two argumens, the first a variable and the second an expression is not trivial.
|
|
735
|
+
Below we create a data frame and we want to write a function that groups data by a variable and
|
|
736
|
+
summarises it by an expression:
|
|
737
|
+
|
|
693
738
|
|
|
694
739
|
``` r
|
|
695
740
|
set.seed(123)
|
|
@@ -704,39 +749,46 @@ df <- data.frame(
|
|
|
704
749
|
as.data.frame(df)
|
|
705
750
|
```
|
|
706
751
|
|
|
707
|
-
|
|
708
|
-
|
|
709
|
-
|
|
710
|
-
|
|
711
|
-
|
|
712
|
-
|
|
752
|
+
```
|
|
753
|
+
## g1 g2 a b
|
|
754
|
+
## 1 1 1 3 3
|
|
755
|
+
## 2 1 2 2 1
|
|
756
|
+
## 3 2 1 5 2
|
|
757
|
+
## 4 2 2 4 5
|
|
758
|
+
## 5 2 1 1 4
|
|
759
|
+
```
|
|
713
760
|
|
|
714
761
|
``` r
|
|
715
762
|
d2 <- df %>%
|
|
716
763
|
group_by(g1) %>%
|
|
717
764
|
summarise(a = mean(a))
|
|
718
|
-
|
|
765
|
+
|
|
719
766
|
as.data.frame(d2)
|
|
720
767
|
```
|
|
721
768
|
|
|
722
|
-
|
|
723
|
-
|
|
724
|
-
|
|
769
|
+
```
|
|
770
|
+
## g1 a
|
|
771
|
+
## 1 1 2.500000
|
|
772
|
+
## 2 2 3.333333
|
|
773
|
+
```
|
|
725
774
|
|
|
726
775
|
``` r
|
|
727
776
|
d2 <- df %>%
|
|
728
777
|
group_by(g2) %>%
|
|
729
778
|
summarise(a = mean(a))
|
|
730
|
-
|
|
731
|
-
as.data.frame(d2)
|
|
779
|
+
|
|
780
|
+
as.data.frame(d2)
|
|
732
781
|
```
|
|
733
782
|
|
|
734
|
-
|
|
735
|
-
|
|
736
|
-
|
|
783
|
+
```
|
|
784
|
+
## g2 a
|
|
785
|
+
## 1 1 3
|
|
786
|
+
## 2 2 3
|
|
787
|
+
```
|
|
737
788
|
|
|
738
789
|
As shown by Hadley, one might expect this function to do the trick:
|
|
739
790
|
|
|
791
|
+
|
|
740
792
|
``` r
|
|
741
793
|
my_summarise <- function(df, group_var) {
|
|
742
794
|
df %>%
|
|
@@ -748,17 +800,15 @@ my_summarise <- function(df, group_var) {
|
|
|
748
800
|
#> Error: Column `group_var` is unknown
|
|
749
801
|
```
|
|
750
802
|
|
|
751
|
-
In order to solve this problem, coding with dplyr requires the
|
|
752
|
-
|
|
753
|
-
|
|
754
|
-
|
|
803
|
+
In order to solve this problem, coding with dplyr requires the introduction of many new concepts
|
|
804
|
+
and functions such as 'quo', 'quos', 'enquo', 'enquos', '!!' (bang bang), '!!!' (triple bang).
|
|
805
|
+
Again, we'll leave to Hadley the explanation on how to use all those functions.
|
|
806
|
+
|
|
807
|
+
Now, let's try to implement the same function in galaaz. The next code block first prints the
|
|
808
|
+
'df' data frame defined previously in R (to access an R variable from Galaaz, we use the tilde
|
|
809
|
+
operator '~' applied to the R variable name as symbol, i.e., ':df'. We then create the
|
|
810
|
+
'my_summarize' method and call it passing the R data frame and the group by variable ':g1':
|
|
755
811
|
|
|
756
|
-
Now, let’s try to implement the same function in galaaz. The next code
|
|
757
|
-
block first prints the ‘df’ data frame defined previously in R (to
|
|
758
|
-
access an R variable from Galaaz, we use the tilde operator ‘~’ applied
|
|
759
|
-
to the R variable name as symbol, i.e., ‘:df’. We then create the
|
|
760
|
-
‘my_summarize’ method and call it passing the R data frame and the group
|
|
761
|
-
by variable ‘:g1’:
|
|
762
812
|
|
|
763
813
|
``` ruby
|
|
764
814
|
puts ~R[:df]
|
|
@@ -773,59 +823,64 @@ end
|
|
|
773
823
|
puts my_summarize(~R[:df], R[:g1])
|
|
774
824
|
```
|
|
775
825
|
|
|
776
|
-
|
|
777
|
-
|
|
778
|
-
|
|
779
|
-
|
|
780
|
-
|
|
781
|
-
|
|
782
|
-
|
|
783
|
-
|
|
784
|
-
|
|
785
|
-
|
|
786
|
-
|
|
787
|
-
|
|
826
|
+
```
|
|
827
|
+
## g1 g2 a b
|
|
828
|
+
## 1 1 1 3 3
|
|
829
|
+
## 2 1 2 2 1
|
|
830
|
+
## 3 2 1 5 2
|
|
831
|
+
## 4 2 2 4 5
|
|
832
|
+
## 5 2 1 1 4
|
|
833
|
+
##
|
|
834
|
+
## # A tibble: 2 × 2
|
|
835
|
+
## g1 a
|
|
836
|
+
## <dbl> <dbl>
|
|
837
|
+
## 1 1 2.5
|
|
838
|
+
## 2 2 3.33
|
|
839
|
+
```
|
|
840
|
+
It works!!! Well, let's make sure this was not just some coincidence
|
|
788
841
|
|
|
789
|
-
It works!!! Well, let’s make sure this was not just some coincidence
|
|
790
842
|
|
|
791
843
|
``` ruby
|
|
792
844
|
puts my_summarize(~R[:df], R[:g2])
|
|
793
845
|
```
|
|
794
846
|
|
|
795
|
-
|
|
796
|
-
|
|
797
|
-
|
|
798
|
-
|
|
799
|
-
|
|
847
|
+
```
|
|
848
|
+
## # A tibble: 2 × 2
|
|
849
|
+
## g2 a
|
|
850
|
+
## <dbl> <dbl>
|
|
851
|
+
## 1 1 3
|
|
852
|
+
## 2 2 3
|
|
853
|
+
```
|
|
800
854
|
|
|
801
|
-
Great, everything is fine! No magic, no new functions, no complexities,
|
|
802
|
-
|
|
803
|
-
certainly feels much safer and easy to implement.
|
|
855
|
+
Great, everything is fine! No magic, no new functions, no complexities, just normal, standard Ruby
|
|
856
|
+
code. If you've ever done NSE in R, this certainly feels much safer and easy to implement.
|
|
804
857
|
|
|
805
858
|
## Different input variables
|
|
806
859
|
|
|
807
|
-
In the previous section we
|
|
808
|
-
for
|
|
809
|
-
|
|
810
|
-
code?
|
|
860
|
+
In the previous section we've managed to get rid of all NSE formulation for a simple example, but
|
|
861
|
+
does this remain true for more complex examples, or will the Galaaz way prove inpractical for
|
|
862
|
+
more complex code?
|
|
811
863
|
|
|
812
|
-
In the next example Hadley proposes us to write a function that given an
|
|
813
|
-
|
|
814
|
-
|
|
864
|
+
In the next example Hadley proposes us to write a function that given an expression such as 'a'
|
|
865
|
+
or 'a * b', calculates three summaries. What we want a function that does the same as these R
|
|
866
|
+
statements:
|
|
815
867
|
|
|
816
|
-
|
|
817
|
-
|
|
818
|
-
|
|
819
|
-
|
|
820
|
-
|
|
868
|
+
```
|
|
869
|
+
summarise(df, mean = mean(a), sum = sum(a), n = n())
|
|
870
|
+
#> # A tibble: 1 x 3
|
|
871
|
+
#> mean sum n
|
|
872
|
+
#> <dbl> <int> <int>
|
|
873
|
+
#> 1 3 15 5
|
|
874
|
+
|
|
875
|
+
summarise(df, mean = mean(a * b), sum = sum(a * b), n = n())
|
|
876
|
+
#> # A tibble: 1 x 3
|
|
877
|
+
#> mean sum n
|
|
878
|
+
#> <dbl> <int> <int>
|
|
879
|
+
#> 1 9 45 5
|
|
880
|
+
```
|
|
821
881
|
|
|
822
|
-
|
|
823
|
-
#> # A tibble: 1 x 3
|
|
824
|
-
#> mean sum n
|
|
825
|
-
#> <dbl> <int> <int>
|
|
826
|
-
#> 1 9 45 5
|
|
882
|
+
Let's try it in galaaz:
|
|
827
883
|
|
|
828
|
-
Let’s try it in galaaz:
|
|
829
884
|
|
|
830
885
|
``` ruby
|
|
831
886
|
def my_summarise2(df, expr)
|
|
@@ -840,49 +895,50 @@ puts my_summarise2((~R[:df]), :a)
|
|
|
840
895
|
puts my_summarise2((~R[:df]), R[:a] * R[:b])
|
|
841
896
|
```
|
|
842
897
|
|
|
843
|
-
|
|
844
|
-
|
|
845
|
-
|
|
846
|
-
|
|
898
|
+
```
|
|
899
|
+
## mean sum n
|
|
900
|
+
## 1 3 15 5
|
|
901
|
+
## mean sum n
|
|
902
|
+
## 1 9 45 5
|
|
903
|
+
```
|
|
847
904
|
|
|
848
|
-
Once again, there is no need to use any special theory or functions.
|
|
849
|
-
|
|
850
|
-
from functions ‘mean’, ‘sum’ and ‘n’.
|
|
905
|
+
Once again, there is no need to use any special theory or functions. The only point to be
|
|
906
|
+
careful about is the use of 'E' to build expressions from functions 'mean', 'sum' and 'n'.
|
|
851
907
|
|
|
852
908
|
## Different input and output variable
|
|
853
909
|
|
|
854
|
-
Now the next challenge presented by Hadley is to vary the name of the
|
|
855
|
-
|
|
856
|
-
|
|
857
|
-
|
|
858
|
-
|
|
859
|
-
|
|
860
|
-
|
|
861
|
-
|
|
862
|
-
|
|
863
|
-
|
|
864
|
-
|
|
865
|
-
|
|
866
|
-
|
|
867
|
-
|
|
868
|
-
|
|
869
|
-
|
|
870
|
-
|
|
871
|
-
|
|
872
|
-
|
|
873
|
-
|
|
874
|
-
|
|
875
|
-
|
|
876
|
-
|
|
877
|
-
|
|
878
|
-
|
|
879
|
-
|
|
880
|
-
In order to solve this problem in R, Hadley needs to introduce some more
|
|
881
|
-
|
|
882
|
-
package ‘rlang’
|
|
910
|
+
Now the next challenge presented by Hadley is to vary the name of the output variables based on
|
|
911
|
+
the received expression. So, if the input expression is 'a', we want our data frame columns to
|
|
912
|
+
be named 'mean\_a' and 'sum\_a'. Now, if the input expression is 'b', columns
|
|
913
|
+
should be named 'mean\_b' and 'sum\_b'.
|
|
914
|
+
|
|
915
|
+
```
|
|
916
|
+
mutate(df, mean_a = mean(a), sum_a = sum(a))
|
|
917
|
+
#> # A tibble: 5 x 6
|
|
918
|
+
#> g1 g2 a b mean_a sum_a
|
|
919
|
+
#> <dbl> <dbl> <int> <int> <dbl> <int>
|
|
920
|
+
#> 1 1 1 1 3 3 15
|
|
921
|
+
#> 2 1 2 4 2 3 15
|
|
922
|
+
#> 3 2 1 2 1 3 15
|
|
923
|
+
#> 4 2 2 5 4 3 15
|
|
924
|
+
#> # … with 1 more row
|
|
925
|
+
|
|
926
|
+
mutate(df, mean_b = mean(b), sum_b = sum(b))
|
|
927
|
+
#> # A tibble: 5 x 6
|
|
928
|
+
#> g1 g2 a b mean_b sum_b
|
|
929
|
+
#> <dbl> <dbl> <int> <int> <dbl> <int>
|
|
930
|
+
#> 1 1 1 1 3 3 15
|
|
931
|
+
#> 2 1 2 4 2 3 15
|
|
932
|
+
#> 3 2 1 2 1 3 15
|
|
933
|
+
#> 4 2 2 5 4 3 15
|
|
934
|
+
#> # … with 1 more row
|
|
935
|
+
```
|
|
936
|
+
In order to solve this problem in R, Hadley needs to introduce some more new functions and notations:
|
|
937
|
+
'quo_name' and the ':=' operator from package 'rlang'
|
|
883
938
|
|
|
884
939
|
Here is our Ruby code:
|
|
885
940
|
|
|
941
|
+
|
|
886
942
|
``` ruby
|
|
887
943
|
def my_mutate(df, expr)
|
|
888
944
|
mean_name = "mean_#{expr.to_s}"
|
|
@@ -896,37 +952,36 @@ puts my_mutate((~R[:df]), :a)
|
|
|
896
952
|
puts my_mutate((~R[:df]), :b)
|
|
897
953
|
```
|
|
898
954
|
|
|
899
|
-
|
|
900
|
-
|
|
901
|
-
|
|
902
|
-
|
|
903
|
-
|
|
904
|
-
|
|
905
|
-
|
|
906
|
-
|
|
907
|
-
|
|
908
|
-
|
|
909
|
-
|
|
910
|
-
|
|
911
|
-
|
|
912
|
-
|
|
913
|
-
|
|
914
|
-
way the arguments to the mutate method were called.
|
|
915
|
-
example we used df.summarise(mean: E.mean(:a),
|
|
916
|
-
|
|
917
|
-
|
|
918
|
-
|
|
919
|
-
|
|
920
|
-
\[explain….\]
|
|
955
|
+
```
|
|
956
|
+
## g1 g2 a b mean_a sum_a
|
|
957
|
+
## 1 1 1 3 3 3 15
|
|
958
|
+
## 2 1 2 2 1 3 15
|
|
959
|
+
## 3 2 1 5 2 3 15
|
|
960
|
+
## 4 2 2 4 5 3 15
|
|
961
|
+
## 5 2 1 1 4 3 15
|
|
962
|
+
## g1 g2 a b mean_b sum_b
|
|
963
|
+
## 1 1 1 3 3 3 15
|
|
964
|
+
## 2 1 2 2 1 3 15
|
|
965
|
+
## 3 2 1 5 2 3 15
|
|
966
|
+
## 4 2 2 4 5 3 15
|
|
967
|
+
## 5 2 1 1 4 3 15
|
|
968
|
+
```
|
|
969
|
+
It really seems that "Non Standard Evaluation" is actually quite standard in Galaaz! But, you
|
|
970
|
+
might have noticed a small change in the way the arguments to the mutate method were called.
|
|
971
|
+
In a previous example we used df.summarise(mean: E.mean(:a), ...) where the column name was
|
|
972
|
+
followed by a ':' colom. In this example, we have df.mutate(mean_name => E.mean(expr), ...)
|
|
973
|
+
and variable mean\_name is not followed by ':' but by '=>'. This is standard Ruby notation.
|
|
974
|
+
|
|
975
|
+
[explain....]
|
|
921
976
|
|
|
922
977
|
## Capturing multiple variables
|
|
923
978
|
|
|
924
|
-
Moving on with new complexities, Hadley proposes us to solve the problem
|
|
925
|
-
|
|
926
|
-
|
|
979
|
+
Moving on with new complexities, Hadley proposes us to solve the problem in which the
|
|
980
|
+
summarise function will receive any number of grouping variables.
|
|
981
|
+
|
|
982
|
+
This again is quite standard Ruby. In order to receive an undefined number of paramenters
|
|
983
|
+
the paramenter is preceded by '*':
|
|
927
984
|
|
|
928
|
-
This again is quite standard Ruby. In order to receive an undefined
|
|
929
|
-
number of paramenters the paramenter is preceded by ’\*’:
|
|
930
985
|
|
|
931
986
|
``` ruby
|
|
932
987
|
def my_summarise3(df, *group_vars)
|
|
@@ -937,89 +992,85 @@ end
|
|
|
937
992
|
puts my_summarise3((~R[:df]), :g1, :g2)
|
|
938
993
|
```
|
|
939
994
|
|
|
940
|
-
|
|
941
|
-
|
|
942
|
-
|
|
943
|
-
|
|
944
|
-
|
|
945
|
-
|
|
946
|
-
|
|
947
|
-
|
|
995
|
+
```
|
|
996
|
+
## # A tibble: 4 × 3
|
|
997
|
+
## # Groups: g1 [2]
|
|
998
|
+
## g1 g2 a
|
|
999
|
+
## <dbl> <dbl> <dbl>
|
|
1000
|
+
## 1 1 1 3
|
|
1001
|
+
## 2 1 2 2
|
|
1002
|
+
## 3 2 1 3
|
|
1003
|
+
## 4 2 2 4
|
|
1004
|
+
```
|
|
948
1005
|
|
|
949
1006
|
# Why does R require NSE and Galaaz does not?
|
|
950
1007
|
|
|
951
|
-
NSE introduces a number of new concepts, such as
|
|
952
|
-
|
|
953
|
-
|
|
954
|
-
|
|
955
|
-
|
|
956
|
-
|
|
957
|
-
|
|
958
|
-
|
|
959
|
-
|
|
960
|
-
|
|
961
|
-
|
|
962
|
-
|
|
963
|
-
|
|
964
|
-
|
|
965
|
-
|
|
966
|
-
|
|
967
|
-
|
|
968
|
-
|
|
969
|
-
|
|
970
|
-
|
|
971
|
-
|
|
972
|
-
|
|
973
|
-
|
|
974
|
-
|
|
975
|
-
|
|
976
|
-
the R function will know how to deal with an input of the form ‘a = b’,
|
|
977
|
-
now for the Ruby developer it might not be immediately clear if it
|
|
978
|
-
should call the function passing the value ‘true’ if variable ‘a’ is
|
|
979
|
-
equal to variable ‘b’ or if it should call the function passing the
|
|
980
|
-
expression ‘R\[:a\].eq R\[:b\]’.
|
|
1008
|
+
NSE introduces a number of new concepts, such as 'quoting', 'quasiquotation', 'unquoting' and
|
|
1009
|
+
'unquote-splicing', while in Galaaz none of those concepts are needed. What gives?
|
|
1010
|
+
|
|
1011
|
+
R is an extremely flexible language and it has lazy evaluation of parameters. When in R a
|
|
1012
|
+
function is called as 'summarise(df, a = b)', the summarise function receives the litteral
|
|
1013
|
+
'a = b' parameter and can work with this as if it were a string. In R, it is not clear what
|
|
1014
|
+
a and b are, they can be expressions or they can be variables, it is up to the function to
|
|
1015
|
+
decide what 'a = b' means.
|
|
1016
|
+
|
|
1017
|
+
In Ruby, there is no lazy evaluation of parameters and 'a' is always a variable and so is 'b'.
|
|
1018
|
+
Variables assume their value as soon as they are used, so 'x = a' is immediately evaluate and
|
|
1019
|
+
variable 'x' will receive the value of variable 'a' as soon as the Ruby statement is executed.
|
|
1020
|
+
Ruby also provides the notion of a symbol; ':a' is a symbol and does not evaluate to anything.
|
|
1021
|
+
Galaaz uses Ruby symbols to build expressions that are not bound to anything: 'R[:a].eq R[:b]' is
|
|
1022
|
+
clearly an expression and has no relationship whatsoever with the statment 'a = b'. By using
|
|
1023
|
+
symbols, variables and expressions all the possible ambiguities that are found in R are
|
|
1024
|
+
eliminated in Galaaz.
|
|
1025
|
+
|
|
1026
|
+
The main problem that remains, is that in R, functions are not clearly documented as what type
|
|
1027
|
+
of input they are expecting, they might be expecting regular variables or they might be
|
|
1028
|
+
expecting expressions and the R function will know how to deal with an input of the form
|
|
1029
|
+
'a = b', now for the Ruby developer it might not be immediately clear if it should call the
|
|
1030
|
+
function passing the value 'true' if variable 'a' is equal to variable 'b' or if it should
|
|
1031
|
+
call the function passing the expression 'R[:a].eq R[:b]'.
|
|
1032
|
+
|
|
981
1033
|
|
|
982
1034
|
# Advanced dplyr features
|
|
983
1035
|
|
|
984
|
-
In the blog: [Programming with dplyr by using
|
|
985
|
-
|
|
986
|
-
Iñaki Úcar shows surprise that some R users are trying to code in dplyr
|
|
987
|
-
avoiding the use of NSE. For instance he says:
|
|
1036
|
+
In the blog: [Programming with dplyr by using dplyr](https://www.r-bloggers.com/programming-with-dplyr-by-using-dplyr/) Iñaki Úcar shows surprise that some R users are trying to code in dplyr avoiding
|
|
1037
|
+
the use of NSE. For instance he says:
|
|
988
1038
|
|
|
989
|
-
> Take the example of seplyr. It stands for standard evaluation dplyr,
|
|
990
|
-
>
|
|
991
|
-
>
|
|
1039
|
+
> Take the example of seplyr. It stands for standard evaluation dplyr, and enables us to
|
|
1040
|
+
> program over dplyr without having “to bring in (or study) any deep-theory or
|
|
1041
|
+
> heavy-weight tools such as rlang/tidyeval”.
|
|
992
1042
|
|
|
993
|
-
For me, there isn
|
|
994
|
-
|
|
995
|
-
to
|
|
996
|
-
|
|
997
|
-
|
|
998
|
-
|
|
1043
|
+
For me, there isn't really any surprise that users are trying to avoid dplyr deep-theory. R
|
|
1044
|
+
users frequently are not programmers and learning to code is already hard business, on top
|
|
1045
|
+
of that, having to learn how to 'quote' or 'enquo' or 'quos' or 'enquos' is not necessarily
|
|
1046
|
+
a 'piece of cake'. So much so, that 'tidyeval' has some more advanced functions that instead
|
|
1047
|
+
of using quoted expressions, uses strings as arguments.
|
|
1048
|
+
|
|
1049
|
+
In the following examples, we show the use of functions 'group\_by\_at', 'summarise\_at' and
|
|
1050
|
+
'rename\_at' that receive strings as argument. The data frame used in 'starwars' that describes
|
|
1051
|
+
features of characters in the Starwars movies:
|
|
999
1052
|
|
|
1000
|
-
In the following examples, we show the use of functions ‘group_by_at’,
|
|
1001
|
-
‘summarise_at’ and ‘rename_at’ that receive strings as argument. The
|
|
1002
|
-
data frame used in ‘starwars’ that describes features of characters in
|
|
1003
|
-
the Starwars movies:
|
|
1004
1053
|
|
|
1005
1054
|
``` ruby
|
|
1006
1055
|
puts (~R[:starwars]).head
|
|
1007
1056
|
```
|
|
1008
1057
|
|
|
1009
|
-
|
|
1010
|
-
|
|
1011
|
-
|
|
1012
|
-
|
|
1013
|
-
|
|
1014
|
-
|
|
1015
|
-
|
|
1016
|
-
|
|
1017
|
-
|
|
1018
|
-
|
|
1019
|
-
|
|
1058
|
+
```
|
|
1059
|
+
## # A tibble: 6 × 14
|
|
1060
|
+
## name height mass hair_color skin_color eye_color birth_year sex
|
|
1061
|
+
## <chr> <int> <dbl> <chr> <chr> <chr> <dbl> <chr>
|
|
1062
|
+
## 1 Luke … 172 77 blond fair blue 19 male
|
|
1063
|
+
## 2 C-3PO 167 75 <NA> gold yellow 112 none
|
|
1064
|
+
## 3 R2-D2 96 32 <NA> white, bl… red 33 none
|
|
1065
|
+
## 4 Darth… 202 136 none white yellow 41.9 male
|
|
1066
|
+
## 5 Leia … 150 49 brown light brown 19 fema…
|
|
1067
|
+
## 6 Owen … 178 120 brown, gr… light blue 52 male
|
|
1068
|
+
## # ℹ 6 more variables: gender <chr>, homeworld <chr>, species <chr>,
|
|
1069
|
+
## # films <list>, vehicles <list>, starships <list>
|
|
1070
|
+
```
|
|
1071
|
+
The grouped_mean function below will receive a grouping variable and calculate summaries for
|
|
1072
|
+
the value\_variables given:
|
|
1020
1073
|
|
|
1021
|
-
The grouped_mean function below will receive a grouping variable and
|
|
1022
|
-
calculate summaries for the value_variables given:
|
|
1023
1074
|
|
|
1024
1075
|
``` r
|
|
1025
1076
|
grouped_mean <- function(data, grouping_variables, value_variables) {
|
|
@@ -1034,40 +1085,45 @@ gm = starwars %>%
|
|
|
1034
1085
|
grouped_mean("eye_color", c("mass", "birth_year"))
|
|
1035
1086
|
```
|
|
1036
1087
|
|
|
1037
|
-
|
|
1038
|
-
|
|
1039
|
-
|
|
1040
|
-
|
|
1041
|
-
|
|
1042
|
-
|
|
1043
|
-
|
|
1044
|
-
|
|
1045
|
-
|
|
1046
|
-
|
|
1088
|
+
```
|
|
1089
|
+
## Warning: `funs()` was deprecated in dplyr 0.8.0.
|
|
1090
|
+
## ℹ Please use a list of either functions or lambdas:
|
|
1091
|
+
##
|
|
1092
|
+
## # Simple named list: list(mean = mean, median = median)
|
|
1093
|
+
##
|
|
1094
|
+
## # Auto named with `tibble::lst()`: tibble::lst(mean, median)
|
|
1095
|
+
##
|
|
1096
|
+
## # Using lambdas list(~ mean(., trim = .2), ~ median(., na.rm = TRUE))
|
|
1097
|
+
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning
|
|
1098
|
+
## was generated.
|
|
1099
|
+
```
|
|
1047
1100
|
|
|
1048
1101
|
``` r
|
|
1049
1102
|
as.data.frame(gm)
|
|
1050
1103
|
```
|
|
1051
1104
|
|
|
1052
|
-
|
|
1053
|
-
|
|
1054
|
-
|
|
1055
|
-
|
|
1056
|
-
|
|
1057
|
-
|
|
1058
|
-
|
|
1059
|
-
|
|
1060
|
-
|
|
1061
|
-
|
|
1062
|
-
|
|
1063
|
-
|
|
1064
|
-
|
|
1065
|
-
|
|
1066
|
-
|
|
1067
|
-
|
|
1105
|
+
```
|
|
1106
|
+
## eye_color mean_mass mean_birth_year count
|
|
1107
|
+
## 1 black 76.28571 33.00000 10
|
|
1108
|
+
## 2 blue 86.51667 67.06923 19
|
|
1109
|
+
## 3 blue-gray 77.00000 57.00000 1
|
|
1110
|
+
## 4 brown 66.09231 108.96429 21
|
|
1111
|
+
## 5 dark NaN NaN 1
|
|
1112
|
+
## 6 gold NaN NaN 1
|
|
1113
|
+
## 7 green, yellow 159.00000 NaN 1
|
|
1114
|
+
## 8 hazel 66.00000 34.50000 3
|
|
1115
|
+
## 9 orange 282.33333 231.00000 8
|
|
1116
|
+
## 10 pink NaN NaN 1
|
|
1117
|
+
## 11 red 81.40000 33.66667 5
|
|
1118
|
+
## 12 red, blue NaN NaN 1
|
|
1119
|
+
## 13 unknown 31.50000 NaN 3
|
|
1120
|
+
## 14 white 48.00000 NaN 1
|
|
1121
|
+
## 15 yellow 81.11111 76.38000 11
|
|
1122
|
+
```
|
|
1068
1123
|
|
|
1069
1124
|
The same code with Galaaz, becomes:
|
|
1070
1125
|
|
|
1126
|
+
|
|
1071
1127
|
``` ruby
|
|
1072
1128
|
def grouped_mean(data, grouping_variables, value_variables)
|
|
1073
1129
|
data.
|
|
@@ -1088,51 +1144,43 @@ puts grouped_mean(
|
|
|
1088
1144
|
E.c("mass", "birth_year"))
|
|
1089
1145
|
```
|
|
1090
1146
|
|
|
1091
|
-
|
|
1092
|
-
|
|
1093
|
-
|
|
1094
|
-
|
|
1095
|
-
|
|
1096
|
-
|
|
1097
|
-
|
|
1098
|
-
|
|
1099
|
-
|
|
1100
|
-
|
|
1101
|
-
|
|
1102
|
-
|
|
1103
|
-
|
|
1104
|
-
|
|
1105
|
-
|
|
1106
|
-
|
|
1107
|
-
|
|
1108
|
-
|
|
1147
|
+
```
|
|
1148
|
+
## # A tibble: 15 × 4
|
|
1149
|
+
## eye_color mean_mass mean_birth_year count
|
|
1150
|
+
## <chr> <dbl> <dbl> <dbl>
|
|
1151
|
+
## 1 black 76.3 33 10
|
|
1152
|
+
## 2 blue 86.5 67.1 19
|
|
1153
|
+
## 3 blue-gray 77 57 1
|
|
1154
|
+
## 4 brown 66.1 109. 21
|
|
1155
|
+
## 5 dark NaN NaN 1
|
|
1156
|
+
## 6 gold NaN NaN 1
|
|
1157
|
+
## 7 green, yellow 159 NaN 1
|
|
1158
|
+
## 8 hazel 66 34.5 3
|
|
1159
|
+
## 9 orange 282. 231 8
|
|
1160
|
+
## 10 pink NaN NaN 1
|
|
1161
|
+
## 11 red 81.4 33.7 5
|
|
1162
|
+
## 12 red, blue NaN NaN 1
|
|
1163
|
+
## 13 unknown 31.5 NaN 3
|
|
1164
|
+
## 14 white 48 NaN 1
|
|
1165
|
+
## 15 yellow 81.1 76.4 11
|
|
1166
|
+
```
|
|
1109
1167
|
|
|
1110
1168
|
# Further reading
|
|
1111
1169
|
|
|
1112
|
-
|
|
1113
|
-
|
|
1114
|
-
|
|
1115
|
-
|
|
1116
|
-
|
|
1117
|
-
|
|
1118
|
-
|
|
1119
|
-
- [How to do reproducible research in Ruby with
|
|
1120
|
-
gKnit](https://towardsdatascience.com/how-to-do-reproducible-research-in-ruby-with-gknit-c26d2684d64e)
|
|
1121
|
-
- [R for Data Science](https://r4ds.had.co.nz/)
|
|
1122
|
-
- [Advanced R](https://adv-r.hadley.nz/)
|
|
1123
|
-
- Historical context: [GraalVM](https://www.graalvm.org/),
|
|
1124
|
-
[TruffleRuby](https://github.com/oracle/truffleruby),
|
|
1125
|
-
[FastR](https://github.com/oracle/fastr)
|
|
1170
|
+
* [JRuby](https://www.jruby.org/) — Ruby on the JVM (Galaaz 2.0)
|
|
1171
|
+
* [How to make Beautiful Ruby Plots with Galaaz](https://medium.freecodecamp.org/how-to-make-beautiful-ruby-plots-with-galaaz-320848058857) (plots; narrative partly pre-2.0)
|
|
1172
|
+
* [Ruby Plotting with Galaaz in GraalVM](https://towardsdatascience.com/ruby-plotting-with-galaaz-an-example-of-tightly-coupling-ruby-and-r-in-graalvm-520b69e21021) (older stack; ideas still useful)
|
|
1173
|
+
* [How to do reproducible research in Ruby with gKnit](https://towardsdatascience.com/how-to-do-reproducible-research-in-ruby-with-gknit-c26d2684d64e)
|
|
1174
|
+
* [R for Data Science](https://r4ds.had.co.nz/)
|
|
1175
|
+
* [Advanced R](https://adv-r.hadley.nz/)
|
|
1176
|
+
* Historical context: [GraalVM](https://www.graalvm.org/), [TruffleRuby](https://github.com/oracle/truffleruby), [FastR](https://github.com/oracle/fastr)
|
|
1126
1177
|
|
|
1127
1178
|
# Conclusion
|
|
1128
1179
|
|
|
1129
|
-
Ruby and Galaaz provide a nice framework for developing code that uses R
|
|
1130
|
-
|
|
1131
|
-
|
|
1132
|
-
|
|
1133
|
-
|
|
1134
|
-
|
|
1135
|
-
|
|
1136
|
-
etc. This simplification comes from the fact that expressions and
|
|
1137
|
-
variables are clearly separated objects, which is not the case in the R
|
|
1138
|
-
language.
|
|
1180
|
+
Ruby and Galaaz provide a nice framework for developing code that uses R functions. Although R is
|
|
1181
|
+
a very powerful and flexible language, sometimes, too much flexibility makes life harder for
|
|
1182
|
+
the casual user. We believe however, that even for the advanced user, Ruby integrated
|
|
1183
|
+
with R throught Galaaz, makes a powerful environment for data analysis. In this blog post we
|
|
1184
|
+
showed how Galaaz consistent syntax eliminates the need for complex constructs such as quoting,
|
|
1185
|
+
enquoting, quasiquotation, etc. This simplification comes from the fact that expressions and
|
|
1186
|
+
variables are clearly separated objects, which is not the case in the R language.
|