galaaz 2.1.7 → 2.1.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (77) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +9 -0
  3. data/blogs/galaaz_ggplot/galaaz_ggplot.Rmd +63 -58
  4. data/blogs/galaaz_ggplot/galaaz_ggplot.log +59 -68
  5. data/blogs/galaaz_ggplot/galaaz_ggplot.md +91 -84
  6. data/blogs/galaaz_ggplot/galaaz_ggplot.tex +125 -94
  7. data/blogs/galaaz_ggplot/galaaz_ggplot_files/figure-html/midwest_rb.png +0 -0
  8. data/blogs/galaaz_ggplot/galaaz_ggplot_files/figure-html/scatter_plot_rb.png +0 -0
  9. data/blogs/galaaz_ggplot/galaaz_ggplot_files/figure-markdown_github/midwest_rb.png +0 -0
  10. data/blogs/galaaz_ggplot/galaaz_ggplot_files/figure-markdown_github/scatter_plot_rb.png +0 -0
  11. data/blogs/gknit/gknit.Rmd +33 -28
  12. data/blogs/gknit/gknit.md +47 -42
  13. data/blogs/gknit/gknit.tex +1368 -0
  14. data/blogs/gknit/gknit_files/figure-html/bubble-1.png +0 -0
  15. data/blogs/gknit/gknit_files/figure-html/diverging_bar.png +0 -0
  16. data/blogs/gknit/gknit_files/figure-latex/bubble-1.png +0 -0
  17. data/blogs/gknit/gknit_files/gknit_files/figure-latex/bubble-1.png +0 -0
  18. data/blogs/manual/manual.Rmd +129 -60
  19. data/blogs/manual/manual.log +289 -545
  20. data/blogs/manual/manual.md +551 -467
  21. data/blogs/manual/manual.tex +1059 -485
  22. data/blogs/manual/manual_files/figure-html/bubble-1.png +0 -0
  23. data/blogs/manual/manual_files/figure-latex/bubble-1.png +0 -0
  24. data/blogs/manual/manual_files/figure-markdown_github/bubble-1.png +0 -0
  25. data/blogs/manual/manual_files/figure-markdown_github/diverging_bar.png +0 -0
  26. data/blogs/manual/manual_files/manual_files/figure-latex/bubble-1.png +0 -0
  27. data/blogs/nse_dplyr/nse_dplyr.Rmd +28 -7
  28. data/blogs/nse_dplyr/nse_dplyr.log +49 -153
  29. data/blogs/nse_dplyr/nse_dplyr.md +676 -705
  30. data/blogs/nse_dplyr/nse_dplyr.tex +1589 -0
  31. data/blogs/oh_my/oh_my.Rmd +193 -55
  32. data/blogs/oh_my/oh_my.log +265 -95
  33. data/blogs/oh_my/oh_my.md +236 -95
  34. data/blogs/oh_my/oh_my.tex +1976 -68
  35. data/blogs/ruby_plot/ruby_plot.Rmd +42 -34
  36. data/blogs/ruby_plot/ruby_plot.log +101 -99
  37. data/blogs/ruby_plot/ruby_plot.md +52 -46
  38. data/blogs/ruby_plot/ruby_plot.tex +134 -102
  39. data/blogs/ruby_plot/ruby_plot_files/figure-html/dose_len.png +0 -0
  40. data/blogs/ruby_plot/ruby_plot_files/figure-html/facet_by_delivery.png +0 -0
  41. data/blogs/ruby_plot/ruby_plot_files/figure-html/facet_by_dose.png +0 -0
  42. data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_by_delivery_color.png +0 -0
  43. data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_by_delivery_color2.png +0 -0
  44. data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_with_decorations.png +0 -0
  45. data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_with_jitter.png +0 -0
  46. data/blogs/ruby_plot/ruby_plot_files/figure-html/facets_with_points.png +0 -0
  47. data/blogs/ruby_plot/ruby_plot_files/figure-html/final_box_plot.png +0 -0
  48. data/blogs/ruby_plot/ruby_plot_files/figure-html/final_violin_plot.png +0 -0
  49. data/blogs/ruby_plot/ruby_plot_files/figure-html/violin_with_jitter.png +0 -0
  50. data/blogs/ruby_plot/ruby_plot_files/figure-latex/dose_len.png +0 -0
  51. data/blogs/ruby_plot/ruby_plot_files/figure-latex/facet_by_delivery.png +0 -0
  52. data/blogs/ruby_plot/ruby_plot_files/figure-latex/facet_by_dose.png +0 -0
  53. data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_by_delivery_color.png +0 -0
  54. data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_by_delivery_color2.png +0 -0
  55. data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_with_decorations.png +0 -0
  56. data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_with_jitter.png +0 -0
  57. data/blogs/ruby_plot/ruby_plot_files/figure-latex/facets_with_points.png +0 -0
  58. data/blogs/ruby_plot/ruby_plot_files/figure-latex/final_box_plot.png +0 -0
  59. data/blogs/ruby_plot/ruby_plot_files/figure-latex/final_violin_plot.png +0 -0
  60. data/blogs/ruby_plot/ruby_plot_files/figure-latex/violin_with_jitter.png +0 -0
  61. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/dose_len.png +0 -0
  62. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facet_by_delivery.png +0 -0
  63. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facet_by_dose.png +0 -0
  64. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_by_delivery_color.png +0 -0
  65. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_by_delivery_color2.png +0 -0
  66. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_with_decorations.png +0 -0
  67. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_with_jitter.png +0 -0
  68. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/facets_with_points.png +0 -0
  69. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/final_box_plot.png +0 -0
  70. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/final_violin_plot.png +0 -0
  71. data/blogs/ruby_plot/ruby_plot_files/ruby_plot_files/figure-latex/violin_with_jitter.png +0 -0
  72. data/lib/galaaz/cli.rb +51 -10
  73. data/script/omarchy/README.md +1 -1
  74. data/script/omarchy/galaaz-guide.sh +1 -1
  75. data/sty/galaaz.sty +22 -0
  76. data/version.rb +1 -1
  77. metadata +19 -1
@@ -1,140 +1,134 @@
1
- ---
2
- title: "Non Standard Evaluation in dplyr with Galaaz"
3
- author:
4
- - "Rodrigo Botafogo"
5
- - "Daniel Mossé - University of Pittsburgh"
6
- tags: [Tech, Data Science, Ruby, R, JRuby, "GNU R", Galaaz, dplyr]
7
- date: "10/05/2019 (narrative updated for Galaaz 2.0, 2026)"
8
- output:
9
- html_document:
10
- self_contained: true
11
- keep_md: true
12
- pdf_document:
13
- includes:
14
- in_header: ["../../sty/galaaz.sty"]
15
- number_sections: yes
16
- toc: true
17
- toc_depth: 2
18
- md_document:
19
- variant: markdown_github
20
- fontsize: 11pt
21
- ---
22
-
23
-
24
-
25
1
  # Introduction
26
2
 
27
- According to Steven Sagaert’s answer on Quora about “Is programming language R overrated?”:
3
+ According to Steven Sagaert’s answer on Quora about “Is programming
4
+ language R overrated?”:
28
5
 
29
- > R is a sophisticated language with an unusual (i.e. non-mainstream) set of features. It‘s
30
- > an impure functional programming language with sophisticated metaprogramming and 3
31
- > different OO systems.
6
+ > R is a sophisticated language with an unusual (i.e. non-mainstream)
7
+ > set of features. It‘s an impure functional programming language with
8
+ > sophisticated metaprogramming and 3 different OO systems.
32
9
 
33
- > Just like common lisp you can completely customise how things work via metaprogramming.
34
- > The biggest example is the tidyverse: by creating it’s own evaluation system (tidyeval)
35
- > was able to create a custom syntax for dplyr.
10
+ > Just like common lisp you can completely customise how things work via
11
+ > metaprogramming. The biggest example is the tidyverse: by creating
12
+ > it’s own evaluation system (tidyeval) was able to create a custom
13
+ > syntax for dplyr.
36
14
 
37
- > Mastering R (the language) and its ecosystem is not a matter of weeks or months but
38
- > takes years. The rabbit hole goes pretty deep…
15
+ > Mastering R (the language) and its ecosystem is not a matter of weeks
16
+ > or months but takes years. The rabbit hole goes pretty deep…
39
17
 
40
- Although a highly configurable language can give programmers a great deal of power,
41
- it can also take years to master—as noted above. Programming with _dplyr_, for instance,
42
- means learning evaluation rules that are not always approachable for **statisticians and
43
- analysts who are not full-time software engineers**. That is not a criticism: R was **built**
44
- for **statisticians** who need trustworthy results on a deadline, not necessarily for building
45
- large applications.
18
+ Although a highly configurable language can give programmers a great
19
+ deal of power, it can also take years to master—as noted above.
20
+ Programming with *dplyr*, for instance, means learning evaluation rules
21
+ that are not always approachable for **statisticians and analysts who
22
+ are not full-time software engineers**. That is not a criticism: R was
23
+ **built** for **statisticians** who need trustworthy results on a
24
+ deadline, not necessarily for building large applications.
46
25
 
47
- **Unfortunately**, when such a user moves on to more **sophisticated** programming patterns,
48
- the learning curve can become a real hurdle.
26
+ **Unfortunately**, when such a user moves on to more **sophisticated**
27
+ programming patterns, the learning curve can become a real hurdle.
49
28
 
50
- In this post we will see how to program with _dplyr_ in Galaaz and how Ruby can simplify
51
- the learning curve of mastering _dplyr_ coding.
29
+ In this post we will see how to program with *dplyr* in Galaaz and how
30
+ Ruby can simplify the learning curve of mastering *dplyr* coding.
52
31
 
53
32
  # But first, what is Galaaz??
54
33
 
55
- Galaaz is a system for tightly coupling Ruby and R. Ruby is a powerful language, with
56
- a large community, a very large set of libraries and great for web development. It is also
57
- easy to learn. However,
58
- it lacks libraries for data science, statistics, scientific plotting and machine learning.
59
- On the other hand, R is considered one of the most powerful languages for solving all of the
60
- above problems. **Python** is a strong competitor, with NumPy, pandas, SciPy, scikit-learn,
61
- and **many thousands** of other packages on PyPI. We will not dwell on R **versus** Python here:
62
- both are excellent languages with different strengths.
63
- Our interest is to bring to yet another excellent language, Ruby, the data science libraries
64
- that it lacks.
65
-
66
- With Galaaz we do not intend to re-implement any of the scientific libraries in R. However, we
67
- allow for very tight coupling between the two languages to the point that the Ruby
68
- developer does not need to know that there is an R engine running. Also, from the point of
69
- view of the R user/developer, Galaaz looks a lot like R, with just minor syntactic difference,
70
- so there is almost no learning curve for the R developer. And as we will see in this
71
- post that programming with _dplyr_ is easier in Galaaz than in R.
72
-
73
- R users are probably quite knowledgeable about _dplyr_. For the Ruby developer, _dplyr_ and
74
- the _tidyverse_ libraries are a set of libraries for data manipulation in R, developed by
75
- Hadley Wickham, Chief Scientist at Posit (formerly RStudio) and a prolific R coder and writer.
76
-
77
- For the coupling of Ruby and R, **Galaaz 2.0** uses **[JRuby](https://www.jruby.org/)** (Ruby on the JVM)
78
- together with **GNU R**. A **bridge** sends expressions and data between Ruby and an R process so that
79
- Ruby can call **dplyr** and the rest of the tidyverse as if they were part of the same workflow.
80
- An **earlier** Galaaz line of work used Oracle’s **GraalVM** with **TruffleRuby** and **FastR** in a single
81
- runtime; that approach is **no longer** the supported stack—see the project **manual** for setup,
82
- **`bin/galaaz-jruby`**, and **gKnit**.
83
-
34
+ Galaaz is a system for tightly coupling Ruby and R. Ruby is a powerful
35
+ language, with a large community, a very large set of libraries and
36
+ great for web development. It is also easy to learn. However, it lacks
37
+ libraries for data science, statistics, scientific plotting and machine
38
+ learning. On the other hand, R is considered one of the most powerful
39
+ languages for solving all of the above problems. **Python** is a strong
40
+ competitor, with NumPy, pandas, SciPy, scikit-learn, and **many
41
+ thousands** of other packages on PyPI. We will not dwell on R **versus**
42
+ Python here: both are excellent languages with different strengths. Our
43
+ interest is to bring to yet another excellent language, Ruby, the data
44
+ science libraries that it lacks.
45
+
46
+ With Galaaz we do not intend to re-implement any of the scientific
47
+ libraries in R. However, we allow for very tight coupling between the
48
+ two languages to the point that the Ruby developer does not need to know
49
+ that there is an R engine running. Also, from the point of view of the R
50
+ user/developer, Galaaz looks a lot like R, with just minor syntactic
51
+ difference, so there is almost no learning curve for the R developer.
52
+ And as we will see in this post that programming with *dplyr* is easier
53
+ in Galaaz than in R.
54
+
55
+ R users are probably quite knowledgeable about *dplyr*. For the Ruby
56
+ developer, *dplyr* and the *tidyverse* libraries are a set of libraries
57
+ for data manipulation in R, developed by Hadley Wickham, Chief Scientist
58
+ at Posit (formerly RStudio) and a prolific R coder and writer.
59
+
60
+ For the coupling of Ruby and R, **Galaaz 2.0** uses
61
+ **[JRuby](https://www.jruby.org/)** (Ruby on the JVM) together with
62
+ **GNU R**. A **bridge** sends expressions and data between Ruby and an R
63
+ process so that Ruby can call **dplyr** and the rest of the tidyverse as
64
+ if they were part of the same workflow. An **earlier** Galaaz line of
65
+ work used Oracle’s **GraalVM** with **TruffleRuby** and **FastR** in a
66
+ single runtime; that approach is **no longer** the supported stack—see
67
+ the project **manual** for setup, **`bin/galaaz-jruby`**, and **gKnit**.
84
68
 
85
69
  # Tidyverse and dplyr
86
70
 
87
- In [What is the tidyverse?](https://rviews.rstudio.com/2017/06/08/what-is-the-tidyverse/) the
88
- tidyverse is explained as follows:
89
-
90
- > The tidyverse is a coherent system of packages for data manipulation, exploration and
91
- > visualization that share a common design philosophy. These were mostly developed by
92
- > Hadley Wickham himself, but they are now being expanded by several contributors. Tidyverse
93
- > packages are intended to make statisticians and data scientists more productive by
94
- > guiding them through workflows that facilitate communication, and result in reproducible
95
- > work products. Fundamentally, the tidyverse is about the connections between the tools
96
- > that make the workflow possible.
97
-
98
- _dplyr_ is one of the many packages that are part of the tidyverse. It is:
99
-
100
- > a grammar of data manipulation, providing a consistent set of verbs that help you solve
101
- > the most common data manipulation challenges:
102
-
103
- > 1. mutate() adds new variables that are functions of existing variables
104
- > 2. select() picks variables based on their names.
105
- > 3. filter() picks cases based on their values.
106
- > 4. summarise() reduces multiple values down to a single summary.
107
- > 5. arrange() changes the ordering of the rows.
108
-
109
- Very often R is used interactively and users use _dplyr_ to manipulate a single dataset
110
- without programming. When users want to replicate their work for
111
- multiple datasets, programming becomes necessary.
71
+ In [What is the
72
+ tidyverse?](https://rviews.rstudio.com/2017/06/08/what-is-the-tidyverse/)
73
+ the tidyverse is explained as follows:
74
+
75
+ > The tidyverse is a coherent system of packages for data manipulation,
76
+ > exploration and visualization that share a common design philosophy.
77
+ > These were mostly developed by Hadley Wickham himself, but they are
78
+ > now being expanded by several contributors. Tidyverse packages are
79
+ > intended to make statisticians and data scientists more productive by
80
+ > guiding them through workflows that facilitate communication, and
81
+ > result in reproducible work products. Fundamentally, the tidyverse is
82
+ > about the connections between the tools that make the workflow
83
+ > possible.
84
+
85
+ *dplyr* is one of the many packages that are part of the tidyverse. It
86
+ is:
87
+
88
+ > a grammar of data manipulation, providing a consistent set of verbs
89
+ > that help you solve the most common data manipulation challenges:
90
+
91
+ > 1. mutate() adds new variables that are functions of existing
92
+ > variables
93
+ > 2. select() picks variables based on their names.
94
+ > 3. filter() picks cases based on their values.
95
+ > 4. summarise() reduces multiple values down to a single summary.
96
+ > 5. arrange() changes the ordering of the rows.
97
+
98
+ Very often R is used interactively and users use *dplyr* to manipulate a
99
+ single dataset without programming. When users want to replicate their
100
+ work for multiple datasets, programming becomes necessary.
112
101
 
113
102
  # Programming with dplyr
114
103
 
115
- In the vignette ["Programming with dplyr"](https://dplyr.tidyverse.org/articles/programming.html),
116
- Hadley Wickham states:
104
+ In the vignette [“Programming with
105
+ dplyr”](https://dplyr.tidyverse.org/articles/programming.html), Hadley
106
+ Wickham states:
117
107
 
118
- > Most dplyr functions use non-standard evaluation (NSE). This is a catch-all term that
119
- > means they don’t follow the usual R rules of evaluation. Instead, they capture the
120
- > expression that you typed and evaluate it in a custom way. This has two main
121
- > benefits for dplyr code:
108
+ > Most dplyr functions use non-standard evaluation (NSE). This is a
109
+ > catch-all term that means they don’t follow the usual R rules of
110
+ > evaluation. Instead, they capture the expression that you typed and
111
+ > evaluate it in a custom way. This has two main benefits for dplyr
112
+ > code:
122
113
 
123
- > Operations on data frames can be expressed succinctly because you don’t need to repeat
124
- > the name of the data frame. For example, you can write filter(df, x == 1, y == 2, z == 3)
125
- > instead of df[df\$x == 1 & df\$y ==2 & df\$z == 3, ].
114
+ > Operations on data frames can be expressed succinctly because you
115
+ > don’t need to repeat the name of the data frame. For example, you can
116
+ > write filter(df, x == 1, y == 2, z == 3) instead of df\[df$x == 1 &
117
+ > df$y ==2 & df$z == 3, \].
126
118
 
127
- > dplyr can choose to compute results in a different way to base R. This is important for
128
- > database backends because dplyr itself doesn’t do any work, but instead generates the SQL
129
- > that tells the database what to do.
119
+ > dplyr can choose to compute results in a different way to base R. This
120
+ > is important for database backends because dplyr itself doesn’t do any
121
+ > work, but instead generates the SQL that tells the database what to
122
+ > do.
130
123
 
131
124
  But then he goes on:
132
125
 
133
- > Unfortunately these benefits do not come for free. There are two main drawbacks:
134
-
135
- > Most dplyr arguments are not referentially transparent. That means you can’t replace a value
136
- > with a seemingly equivalent object that you’ve defined elsewhere. In other words, this code:
126
+ > Unfortunately these benefits do not come for free. There are two main
127
+ > drawbacks:
137
128
 
129
+ > Most dplyr arguments are not referentially transparent. That means you
130
+ > can’t replace a value with a seemingly equivalent object that you’ve
131
+ > defined elsewhere. In other words, this code:
138
132
 
139
133
  ``` r
140
134
  df <- data.frame(x = 1:3, y = 3:1)
@@ -144,8 +138,8 @@ print(filter(df, x == 1))
144
138
  #> <int> <int>
145
139
  #> 1 1 3
146
140
  ```
147
- > Is not equivalent to this code:
148
141
 
142
+ > Is not equivalent to this code:
149
143
 
150
144
  ``` r
151
145
  my_var <- x
@@ -153,109 +147,110 @@ my_var <- x
153
147
  filter(df, my_var == 1)
154
148
  #> Error: object 'my_var' not found
155
149
  ```
156
- > This makes it hard to create functions with arguments that change how dplyr verbs are computed.
157
150
 
158
- As a result of this, programming with _dplyr_ requires learning a set of new ideas and concepts.
159
- In this vignette Hadley goes on showing how to program ever more difficult problems with _dplyr_,
160
- showing the problems it faces and the new concepts needed to solve them.
151
+ > This makes it hard to create functions with arguments that change how
152
+ > dplyr verbs are computed.
153
+
154
+ As a result of this, programming with *dplyr* requires learning a set of
155
+ new ideas and concepts. In this vignette Hadley goes on showing how to
156
+ program ever more difficult problems with *dplyr*, showing the problems
157
+ it faces and the new concepts needed to solve them.
161
158
 
162
- In this blog, we will look at all the problems presented by Harley on the vignette and show how
163
- those same problems can be solved using Galaaz and the Ruby language.
159
+ In this blog, we will look at all the problems presented by Harley on
160
+ the vignette and show how those same problems can be solved using Galaaz
161
+ and the Ruby language.
164
162
 
165
- This blog is organized as follows: first we show how to write expressions using Galaaz.
166
- Expressions are a fundamental concept in _dplyr_ and are not part of basic Ruby. We extend
167
- the Ruby language create a manipulate expressions that will be used by _dplyr_ functions.
163
+ This blog is organized as follows: first we show how to write
164
+ expressions using Galaaz.
165
+ Expressions are a fundamental concept in *dplyr* and are not part of
166
+ basic Ruby. We extend the Ruby language create a manipulate expressions
167
+ that will be used by *dplyr* functions.
168
168
 
169
- Then we show very succintly how Ruby and R can be integrated and how R functions are
170
- transparently called from Ruby. Galaaz [user manual](https://github.com/rbotafogo/galaaz/wiki)
171
- (still in development) goes in much deeper detail about this integration.
169
+ Then we show very succintly how Ruby and R can be integrated and how R
170
+ functions are transparently called from Ruby. Galaaz [user
171
+ manual](https://github.com/rbotafogo/galaaz/wiki) (still in development)
172
+ goes in much deeper detail about this integration.
172
173
 
173
- Next in section "Data manipulation wiht _dplyr_" we go through all the problems on the
174
- _dplyr_ vignette and look at how they are solved in Galaaz. We then discuss why programming
175
- with Galaaz and _dplyr_ is easier than programming with _dplyr_ in plain R.
174
+ Next in section “Data manipulation wiht *dplyr*” we go through all the
175
+ problems on the *dplyr* vignette and look at how they are solved in
176
+ Galaaz. We then discuss why programming with Galaaz and *dplyr* is
177
+ easier than programming with *dplyr* in plain R.
176
178
 
177
- The following section looks at another more advanced problem and shows that Galaaz can still
178
- handle it without any difficulty. We then provide further reading and concluding remarks.
179
+ The following section looks at another more advanced problem and shows
180
+ that Galaaz can still handle it without any difficulty. We then provide
181
+ further reading and concluding remarks.
179
182
 
180
183
  # Writing Expressions in Galaaz
181
184
 
182
- Galaaz extends Ruby to work with expressions, similar to R's expressions build with 'quote'
183
- (base R) or 'quo' (tidyverse). Expressions in this context are like mathematical expressions or
184
- formulae. For instance, in mathematics, the expression $y = sin(x)$ describes a function but cannot
185
- be computed unless the value of $x$ is bound to some value.
185
+ Galaaz extends Ruby to work with expressions, similar to R’s expressions
186
+ build with ‘quote’ (base R) or ‘quo’ (tidyverse). Expressions in this
187
+ context are like mathematical expressions or formulae. For instance, in
188
+ mathematics, the expression *y* = *s**i**n*(*x*) describes a function
189
+ but cannot be computed unless the value of *x* is bound to some value.
186
190
 
187
- Expressions are fundamental in _dplyr_ programming as they are the input to _dplyr_ functions,
188
- for instance, as we will see shortly, if a data frame has a column named 'x' and we want
189
- to add another column, y, to this dataframe that has the values of 'x' times 2, then we would
190
- call a _dplyr_ function with the expression 'y = x * 2'.
191
+ Expressions are fundamental in *dplyr* programming as they are the input
192
+ to *dplyr* functions, for instance, as we will see shortly, if a data
193
+ frame has a column named ‘x’ and we want to add another column, y, to
194
+ this dataframe that has the values of ‘x’ times 2, then we would call a
195
+ *dplyr* function with the expression ‘y = x \* 2’.
191
196
 
192
197
  ## A note on notation
193
198
 
194
- This blog was written in Rmarkdown and automatically converted to HTML or PDF (depending on
195
- where you are reading this blog) with gKnit (a tool provided by Galaaz). In Rmarkdown, it is
196
- possible to write text and code blocks that are executed to generate the final report. Code
197
- blocks appear inside a 'box' and the result of their execution appear either in another type
198
- of 'box' with a different background (HTML) or as normal text (PDF). Every output line from
199
- the code execution is preceded by '##'.
199
+ This blog was written in Rmarkdown and automatically converted to HTML
200
+ or PDF (depending on where you are reading this blog) with gKnit (a tool
201
+ provided by Galaaz). In Rmarkdown, it is possible to write text and code
202
+ blocks that are executed to generate the final report. Code blocks
203
+ appear inside a ‘box’ and the result of their execution appear either in
204
+ another type of ‘box’ with a different background (HTML) or as normal
205
+ text (PDF). Every output line from the code execution is preceded by
206
+ ‘\##’.
200
207
 
201
208
  ## Expressions from operators
202
209
 
203
- The code below creates an expression summing two symbols. Note that :a and :b are Ruby symbols and
204
- are not bound to any values at the time of expression definition:
205
-
210
+ The code below creates an expression summing two symbols. Note that :a
211
+ and :b are Ruby symbols and are not bound to any values at the time of
212
+ expression definition:
206
213
 
207
214
  ``` ruby
208
- exp1 = :a + :b
209
- puts exp1
215
+ begin
216
+ exp1 = :a + :b
217
+ puts exp1
218
+ rescue => e
219
+ # Bare Symbol#+ is not expression sugar in Galaaz 2.0; short error for PDF.
220
+ puts "#{e.class}: #{e.message}"
221
+ end
210
222
  ```
211
223
 
212
- ```
213
- ## undefined method '+' for an instance of Symbol
214
- ```
224
+ ## NoMethodError: undefined method '+' for an instance of Symbol
215
225
 
216
- ```
217
- ## /home/rbotafogo/desenv_linux/galaaz/lib/util/exec_ruby.rb:170:in 'exec_ruby'
218
- ## org/jruby/RubyKernel.java:1268:in 'eval'
219
- ## /home/rbotafogo/desenv_linux/galaaz/lib/util/exec_ruby.rb:169:in 'exec_ruby'
220
- ## /home/rbotafogo/desenv_linux/galaaz/lib/gknit/knitr_engine.rb:777:in 'block in initialize'
221
- ## org/jruby/RubyBasicObject.java:2695:in 'instance_eval'
222
- ## org/jruby/RubyBasicObject.java:2723:in 'instance_eval'
223
- ## /home/rbotafogo/desenv_linux/galaaz/lib/gknit/knitr_engine.rb:748:in 'block in initialize'
224
- ## /home/rbotafogo/desenv_linux/galaaz/lib/R_interface/new_bridge_adapter.rb:358:in 'block in register_callback_proc_stub'
225
- ## /home/rbotafogo/desenv_linux/galaaz/lib/new_bridge/session_client.rb:413:in 'block in handle_call'
226
- ```
227
226
  In Galaaz, we can build any complex mathematical expression such as:
228
227
 
229
-
230
228
  ``` ruby
231
229
  exp2 = (R[:a] + R[:b]) * 2.0 + R[:c] ** 2 / R[:z]
232
230
  puts exp2
233
231
  ```
234
232
 
235
- ```
236
- ## a + b * 2.0 + c ^ 2L / z
237
- ```
238
- Expressions are printed with the same format as the equivalent R expressions. The 'L' after
239
- 2 indicates that 2 is an integer.
233
+ ## a + b * 2.0 + c ^ 2L / z
240
234
 
241
- The R developer should note that in R, if she writes the
242
- number '2', the R interpreter will convert it to float. In order to get an interger she
243
- should write '2L'. Galaaz follows Ruby notation and '2' is an integer, while '2.0' is a
244
- float.
235
+ Expressions are printed with the same format as the equivalent R
236
+ expressions. The ‘L’ after 2 indicates that 2 is an integer.
245
237
 
246
- It is also possible to use inequality operators in building expressions:
238
+ The R developer should note that in R, if she writes the number ‘2’, the
239
+ R interpreter will convert it to float. In order to get an interger she
240
+ should write ‘2L’. Galaaz follows Ruby notation and ‘2’ is an integer,
241
+ while ‘2.0’ is a float.
247
242
 
243
+ It is also possible to use inequality operators in building expressions:
248
244
 
249
245
  ``` ruby
250
246
  exp3 = (R[:a] + R[:b]) >= R[:z]
251
247
  puts exp3
252
248
  ```
253
249
 
254
- ```
255
- ## a + b >= z
256
- ```
257
- Expressions' definition can also make use of normal Ruby variables without any problem:
250
+ ## a + b >= z
258
251
 
252
+ Expressions’ definition can also make use of normal Ruby variables
253
+ without any problem:
259
254
 
260
255
  ``` ruby
261
256
  x = 20
@@ -264,133 +259,116 @@ exp_var = (R[:a] + R[:b]) * x <= R[:z] - y
264
259
  puts exp_var
265
260
  ```
266
261
 
267
- ```
268
- ## a + b * 20L <= z - 30.0
269
- ```
270
-
271
- Galaaz provides both symbolic representations for operators, such as (>, <, !=) as functional
272
- notation for those operators such as (.gt, .ge, etc.). So the same expression written
273
- above can also be written as
262
+ ## a + b * 20L <= z - 30.0
274
263
 
264
+ Galaaz provides both symbolic representations for operators, such as
265
+ (\>, \<, !=) as functional notation for those operators such as (.gt,
266
+ .ge, etc.). So the same expression written above can also be written as
275
267
 
276
268
  ``` ruby
277
269
  exp4 = (R[:a] + R[:b]).ge R[:z]
278
270
  puts exp4
279
271
  ```
280
272
 
281
- ```
282
- ## a + b >= z
283
- ```
273
+ ## a + b >= z
284
274
 
285
- Two types of expressions, however, can only be created with the functional representation
286
- of the operators. Those are expressions involving '==', and '='. This is the case since
287
- those symbols have special meaning in Ruby and should not be redefined.
288
-
289
- In order to write an expression involving '==' we
290
- need to use the method '.eq' and for '=' we need the function '.assign':
275
+ Two types of expressions, however, can only be created with the
276
+ functional representation of the operators. Those are expressions
277
+ involving ‘==’, and ‘=’. This is the case since those symbols have
278
+ special meaning in Ruby and should not be redefined.
291
279
 
280
+ In order to write an expression involving ‘==’ we need to use the method
281
+ ‘.eq’ and for ‘=’ we need the function ‘.assign’:
292
282
 
293
283
  ``` ruby
294
284
  exp5 = (R[:a] + R[:b]).eq R[:z]
295
285
  puts exp5
296
286
  ```
297
287
 
298
- ```
299
- ## a + b == z
300
- ```
301
-
288
+ ## a + b == z
302
289
 
303
290
  ``` ruby
304
291
  exp6 = R[:y].assign R[:a] + R[:b]
305
292
  puts exp6
306
293
  ```
307
294
 
308
- ```
309
- ## y <- a + b
310
- ```
311
- Users should be careful when writing expressions not to inadvertently use '==' or '=' as
312
- this will generate an error, that might be a bit cryptic (in future releases of Galaza, we
313
- plan to improve the error message).
295
+ ## y <- a + b
314
296
 
297
+ Users should be careful when writing expressions not to inadvertently
298
+ use ‘==’ or ‘=’ as this will generate an error, that might be a bit
299
+ cryptic (in future releases of Galaza, we plan to improve the error
300
+ message).
315
301
 
316
302
  ``` ruby
317
303
  exp_wrong = (R[:a] + R[:b]) == R[:z]
318
304
  puts exp_wrong
319
305
  ```
320
306
 
321
- ```
322
- ## false
323
- ```
324
- The problem lies with the fact that
325
- when using '==' we are comparing expression (R[:a] + R[:b]) to expression R[:z] with '=='. When this
326
- comparison is executed, the system tries to evaluate :a, :b and :z, and those symbols, at
327
- this time, are not bound to anything giving the "object 'a' not found" message.
307
+ ## false
328
308
 
329
- ## Expressions with R methods
309
+ The problem lies with the fact that when using ‘==’ we are comparing
310
+ expression (R\[:a\] + R\[:b\]) to expression R\[:z\] with ‘==’. When
311
+ this comparison is executed, the system tries to evaluate :a, :b and :z,
312
+ and those symbols, at this time, are not bound to anything giving the
313
+ “object ‘a’ not found” message.
330
314
 
331
- It is often necessary to create an expression that uses a method or function. For instance, in
332
- mathematics, it's quite natural to write an expressin such as $y = sin(x)$. In this case, the
333
- 'sin' function is part of the expression and should not be immediately executed. When we want
334
- the function to be part of the expression, we call the function preceeding it
335
- by the letter E, such as 'E.sin(x)'
315
+ ## Expressions with R methods
336
316
 
317
+ It is often necessary to create an expression that uses a method or
318
+ function. For instance, in mathematics, it’s quite natural to write an
319
+ expressin such as *y* = *s**i**n*(*x*). In this case, the ‘sin’ function
320
+ is part of the expression and should not be immediately executed. When
321
+ we want the function to be part of the expression, we call the function
322
+ preceeding it by the letter E, such as ‘E.sin(x)’
337
323
 
338
324
  ``` ruby
339
325
  exp7 = R[:y].assign E.sin(R[:x])
340
326
  puts exp7
341
327
  ```
342
328
 
343
- ```
344
- ## y <- sin(x)
345
- ```
346
- Function expressions can also be written using '.' notation:
329
+ ## y <- sin(x)
347
330
 
331
+ Function expressions can also be written using ‘.’ notation:
348
332
 
349
333
  ``` ruby
350
334
  exp8 = R[:y].assign R[:x].sin
351
335
  puts exp8
352
336
  ```
353
337
 
354
- ```
355
- ## y <- sin(x)
356
- ```
357
- When a function has multiple arguments, the first one can be used before the '.'. For instance,
358
- the R concatenate function 'c', that concatenates two or more arguments can be part of
359
- an expression as:
338
+ ## y <- sin(x)
360
339
 
340
+ When a function has multiple arguments, the first one can be used before
341
+ the ‘.’. For instance, the R concatenate function ‘c’, that concatenates
342
+ two or more arguments can be part of an expression as:
361
343
 
362
344
  ``` ruby
363
345
  exp9 = R[:x].c(R[:y])
364
346
  puts exp9
365
347
  ```
366
348
 
367
- ```
368
- ## c(x, y)
369
- ```
370
- Note that this gives an OO feeling to the code, as if we were saying 'x' concatenates 'y'. As a
371
- side note, '.' notation can be used as the R pipe operator '%>%', but is more general than the
372
- pipe.
349
+ ## c(x, y)
350
+
351
+ Note that this gives an OO feeling to the code, as if we were saying ‘x’
352
+ concatenates ‘y’. As a side note, ‘.’ notation can be used as the R pipe
353
+ operator ‘%\>%’, but is more general than the pipe.
373
354
 
374
355
  ## Evaluating an Expression
375
356
 
376
- Although we are mainly focusing on expressions to pass them to _dplyr_ functions, expressions
377
- can be evaluated by calling function 'eval' with a binding.
357
+ Although we are mainly focusing on expressions to pass them to *dplyr*
358
+ functions, expressions can be evaluated by calling function ‘eval’ with
359
+ a binding.
378
360
 
379
361
  A binding can be provided with a list or a data frame as shown below:
380
362
 
381
-
382
363
  ``` ruby
383
364
  exp = (R[:a] + R[:b]) * 2.0 + R[:c] ** 2 / R[:z]
384
365
  puts exp.eval(R.list(a: 10, b: 20, c: 30, z: 40))
385
366
  ```
386
367
 
387
- ```
388
- ## [1] 72.5
389
- ```
368
+ ## [1] 72.5
390
369
 
391
370
  with a data frame:
392
371
 
393
-
394
372
  ``` ruby
395
373
  df = R.data__frame(
396
374
  a: R.c(1, 2, 3),
@@ -401,90 +379,83 @@ df = R.data__frame(
401
379
  puts exp.eval(df)
402
380
  ```
403
381
 
404
- ```
405
- ## [1] 31 62 93
406
- ```
382
+ ## [1] 31 62 93
407
383
 
408
384
  # Using Galaaz to call R functions
409
385
 
410
- Galaaz tries to emulate as closely as possible the way R functions are called and migrating from
411
- R to Galaaz should be quite easy requiring only minor syntactic changes to an R script. In
412
- this post, we do not have enough space to write a complete manual on Galaaz
413
- (a short manual can be found at: https://www.rubydoc.info/gems/galaaz/0.4.9), so we will
414
- present only a few examples scripts using Galaaz.
415
-
416
- Basically, to call an R function from Ruby with Galaaz, one only needs to preced the function
417
- with 'R.'. For instance, to create a vector in R, the 'c' function is used. In Galaaz, a
418
- vector can be created by using 'R.c':
386
+ Galaaz tries to emulate as closely as possible the way R functions are
387
+ called and migrating from R to Galaaz should be quite easy requiring
388
+ only minor syntactic changes to an R script. In this post, we do not
389
+ have enough space to write a complete manual on Galaaz (a short manual
390
+ can be found at: <https://www.rubydoc.info/gems/galaaz/0.4.9>), so we
391
+ will present only a few examples scripts using Galaaz.
419
392
 
393
+ Basically, to call an R function from Ruby with Galaaz, one only needs
394
+ to preced the function with ‘R.’. For instance, to create a vector in R,
395
+ the ‘c’ function is used. In Galaaz, a vector can be created by using
396
+ ‘R.c’:
420
397
 
421
398
  ``` ruby
422
399
  vec = R.c(1.0, 2, 3)
423
400
  puts vec
424
401
  ```
425
402
 
426
- ```
427
- ## [1] 1 2 3
428
- ```
429
- A list is created in R with the 'list' function, so in Galaaz we do:
403
+ ## [1] 1 2 3
430
404
 
405
+ A list is created in R with the ‘list’ function, so in Galaaz we do:
431
406
 
432
407
  ``` ruby
433
408
  list = R.list(a: 1.0, b: 2, c: 3)
434
409
  puts list
435
410
  ```
436
411
 
437
- ```
438
- ## $a
439
- ## [1] 1
440
- ##
441
- ## $b
442
- ## [1] 2
443
- ##
444
- ## $c
445
- ## [1] 3
446
- ```
447
- Note that we can use named arguments in our list. The same code in R would be:
412
+ ## $a
413
+ ## [1] 1
414
+ ##
415
+ ## $b
416
+ ## [1] 2
417
+ ##
418
+ ## $c
419
+ ## [1] 3
448
420
 
421
+ Note that we can use named arguments in our list. The same code in R
422
+ would be:
449
423
 
450
424
  ``` r
451
425
  lst = list(a = 1, b = 2L, c = 3L)
452
426
  print(lst)
453
427
  ```
454
428
 
455
- ```
456
- ## $a
457
- ## [1] 1
458
- ##
459
- ## $b
460
- ## [1] 2
461
- ##
462
- ## $c
463
- ## [1] 3
464
- ```
465
- Now, let's say that 'x' is an angle of 45$^\circ$ and we acttually want to create
466
- the expression $y = sin(45^\circ)$, which is $y = 0.850...$. In this case,
467
- we will use 'R.sin':
429
+ ## $a
430
+ ## [1] 1
431
+ ##
432
+ ## $b
433
+ ## [1] 2
434
+ ##
435
+ ## $c
436
+ ## [1] 3
468
437
 
438
+ Now, let’s say that ‘x’ is an angle of 45<sup>∘</sup> and we acttually
439
+ want to create the expression *y* = *s**i**n*(45<sup>∘</sup>), which is
440
+ *y* = 0.850.... In this case, we will use ‘R.sin’:
469
441
 
470
442
  ``` ruby
471
443
  exp10 = R[:y].assign R.sin(45)
472
444
  puts exp10
473
445
  ```
474
446
 
475
- ```
476
- ## y <- 0.850903524534118
477
- ```
478
-
479
- # Data manipulation wiht _dplyr_
447
+ ## y <- 0.850903524534118
480
448
 
481
- In this section we will give a brief tour _dplyr_'s usage in Galaaz and how to manipulate
482
- data in Ruby with it. This section will follow [_dplyr_'s vignette](https://dplyr.tidyverse.org/articles/dplyr.html) that explores the nycflights13 data set. This dataset contains all 336776
483
- flights that departed from New York City in 2013. The data comes from the US Bureau of
484
- Transportation Statistics.
449
+ # Data manipulation wiht *dplyr*
485
450
 
486
- Let's start by taking a look at this dataset:
451
+ In this section we will give a brief tour *dplyr*’s usage in Galaaz and
452
+ how to manipulate data in Ruby with it. This section will follow
453
+ [*dplyr*’s vignette](https://dplyr.tidyverse.org/articles/dplyr.html)
454
+ that explores the nycflights13 data set. This dataset contains all
455
+ 336776 flights that departed from New York City in 2013. The data comes
456
+ from the US Bureau of Transportation Statistics.
487
457
 
458
+ Let’s start by taking a look at this dataset:
488
459
 
489
460
  ``` ruby
490
461
  R.library('nycflights13')
@@ -494,140 +465,137 @@ puts ~R[:flights].dim
494
465
  ~R[:flights].str
495
466
  ```
496
467
 
497
- ```
498
- ## ~(dim(flights))
499
- ## <environment: 0x5f0b6a5b7790>
500
- ```
501
-
502
- Now, let's use a first verb of _dplyr_: 'filter'. This verb, obviously, will filter the data
503
- by the given expression. In the next block, we filter by columns 'month' and 'day'. The
504
- first argument to the filter function is symbol ':flights'. A Ruby symbol, when given to
505
- an R function will convert to the R variable of the same name, in this case 'flights', that
506
- holds the nycflights13 data frame.
468
+ ## ~(dim(flights))
469
+ ## <environment: 0x57d8ad019298>
507
470
 
508
- The second and third arguments are expressions that will be used by the filter function to
509
- filter by columns, looking for entries in which the month and day are equal to 1.
471
+ Now, let’s use a first verb of *dplyr*: ‘filter’. This verb, obviously,
472
+ will filter the data by the given expression. In the next block, we
473
+ filter by columns ‘month’ and ‘day’. The first argument to the filter
474
+ function is symbol ‘:flights’. A Ruby symbol, when given to an R
475
+ function will convert to the R variable of the same name, in this case
476
+ ‘flights’, that holds the nycflights13 data frame.
510
477
 
478
+ The second and third arguments are expressions that will be used by the
479
+ filter function to filter by columns, looking for entries in which the
480
+ month and day are equal to 1.
511
481
 
512
482
  ``` ruby
513
483
  puts R.filter(:flights, (R[:month].eq 1), (R[:day].eq 1))
514
484
  ```
515
485
 
516
- ```
517
- ## # A tibble: 842 × 19
518
- ## year month day dep_time sched_dep_time dep_delay arr_time sched_arr_time
519
- ## <int> <int> <int> <int> <int> <dbl> <int> <int>
520
- ## 1 2013 1 1 517 515 2 830 819
521
- ## 2 2013 1 1 533 529 4 850 830
522
- ## 3 2013 1 1 542 540 2 923 850
523
- ## 4 2013 1 1 544 545 -1 1004 1022
524
- ## 5 2013 1 1 554 600 -6 812 837
525
- ## 6 2013 1 1 554 558 -4 740 728
526
- ## 7 2013 1 1 555 600 -5 913 854
527
- ## 8 2013 1 1 557 600 -3 709 723
528
- ## 9 2013 1 1 557 600 -3 838 846
529
- ## 10 2013 1 1 558 600 -2 753 745
530
- ## # ℹ 832 more rows
531
- ## # ℹ 11 more variables: arr_delay <dbl>, carrier <chr>, flight <int>,
532
- ## # tailnum <chr>, origin <chr>, dest <chr>, air_time <dbl>, distance <dbl>,
533
- ## # hour <dbl>, minute <dbl>, time_hour <dttm>
534
- ```
535
-
536
-
537
- ## Programming with _dplyr_: problems and how to solve them in Galaaz
538
-
539
- In this section we look at the list of problems that Hadley describes in the "Programming with dplyr"
540
- vignette and show how those problems are solved and coded with Galaaz. Readers interested in
541
- how those problems are treated in _dplyr_ should read the vignette and use it as a comparison with
542
- this blog.
486
+ ## # A tibble: 842 × 19
487
+ ## year month day dep_time sched_dep_time dep_delay arr_time
488
+ ## <int> <int> <int> <int> <int> <dbl> <int>
489
+ ## 1 2013 1 1 517 515 2 830
490
+ ## 2 2013 1 1 533 529 4 850
491
+ ## 3 2013 1 1 542 540 2 923
492
+ ## 4 2013 1 1 544 545 -1 1004
493
+ ## 5 2013 1 1 554 600 -6 812
494
+ ## 6 2013 1 1 554 558 -4 740
495
+ ## 7 2013 1 1 555 600 -5 913
496
+ ## 8 2013 1 1 557 600 -3 709
497
+ ## 9 2013 1 1 557 600 -3 838
498
+ ## 10 2013 1 1 558 600 -2 753
499
+ ## # ℹ 832 more rows
500
+ ## # ℹ 12 more variables: sched_arr_time <int>, arr_delay <dbl>,
501
+ ## # carrier <chr>, flight <int>, tailnum <chr>, origin <chr>,
502
+ ## # dest <chr>, air_time <dbl>, distance <dbl>, hour <dbl>,
503
+ ## # minute <dbl>, time_hour <dttm>
504
+
505
+ ## Programming with *dplyr*: problems and how to solve them in Galaaz
506
+
507
+ In this section we look at the list of problems that Hadley describes in
508
+ the “Programming with dplyr” vignette and show how those problems are
509
+ solved and coded with Galaaz. Readers interested in how those problems
510
+ are treated in *dplyr* should read the vignette and use it as a
511
+ comparison with this blog.
543
512
 
544
513
  ## Filtering using expressions
545
514
 
546
- Now that we know how to write expressions and call R functions, let's do some data manipulation in
547
- Galaaz. Let's first start by creating a data frame. In R, the 'data.frame' function creates a
548
- data frame. In Ruby, writing 'data.frame' will not parse as a single object. To call R
549
- functions that have a '.' in them, we need to substitute the '.' with '__'. So, method
550
- 'data.frame' in R, is called in Galaaz as 'R.data\_\_frame':
551
-
515
+ Now that we know how to write expressions and call R functions, let’s do
516
+ some data manipulation in Galaaz. Let’s first start by creating a data
517
+ frame. In R, the ‘data.frame’ function creates a data frame. In Ruby,
518
+ writing ‘data.frame’ will not parse as a single object. To call R
519
+ functions that have a ‘.’ in them, we need to substitute the ‘.’ with
520
+ ’\_\_‘. So, method ’data.frame’ in R, is called in Galaaz as
521
+ ‘R.data\_\_frame’:
552
522
 
553
523
  ``` ruby
554
524
  df = R.data__frame(x: (1..3), y: (3..1))
555
525
  puts df
556
526
  ```
557
527
 
558
- ```
559
- ## x y
560
- ## 1 1 3
561
- ## 2 2 2
562
- ## 3 3 1
563
- ```
564
-
565
- _dplyr_ provides the 'filter' function, that filters data in a data brame. The 'filter'
566
- function can be called on this data frame either by using 'R.filter(df, ...)' or
567
- by using dot notation.
528
+ ## x y
529
+ ## 1 1 3
530
+ ## 2 2 2
531
+ ## 3 3 1
568
532
 
569
- -------FIX---------
533
+ *dplyr* provides the ‘filter’ function, that filters data in a data
534
+ brame. The ‘filter’ function can be called on this data frame either by
535
+ using ‘R.filter(df, …)’ or by using dot notation.
570
536
 
571
- We prefer to use dot notation as shown below. The argument to 'filter' should be an
572
- expression. Note that if we gave to filter a Ruby expression such as
573
- 'x == 1', we would get an error, since there is no variable 'x' defined and if 'x' was a variable
574
- then 'x == 1' would either be 'true' or 'false'. Our goal is to filter our data frame returning
575
- all rows in which the 'x' value is equal to 1. To express this we want: 'R[:x].eq 1', where :x will
576
- be interpreted by filter as the 'x' column.
537
+ ——-FIX———
577
538
 
539
+ We prefer to use dot notation as shown below. The argument to ‘filter’
540
+ should be an expression. Note that if we gave to filter a Ruby
541
+ expression such as ‘x == 1’, we would get an error, since there is no
542
+ variable ‘x’ defined and if ‘x’ was a variable then ‘x == 1’ would
543
+ either be ‘true’ or ‘false’. Our goal is to filter our data frame
544
+ returning all rows in which the ‘x’ value is equal to 1. To express this
545
+ we want: ‘R\[:x\].eq 1’, where :x will be interpreted by filter as the
546
+ ‘x’ column.
578
547
 
579
548
  ``` ruby
580
549
  puts df.filter(R[:x].eq 1)
581
550
  ```
582
551
 
583
- ```
584
- ## x y
585
- ## 1 1 3
586
- ```
587
- In R, and when coding with 'tidyverse', arguments to a function are usually not
588
- *referencially transparent*. That is, you can’t replace a value with a seemingly equivalent
589
- object that you’ve defined elsewhere. In other words, this code
552
+ ## x y
553
+ ## 1 1 3
590
554
 
555
+ In R, and when coding with ‘tidyverse’, arguments to a function are
556
+ usually not *referencially transparent*. That is, you can’t replace a
557
+ value with a seemingly equivalent object that you’ve defined elsewhere.
558
+ In other words, this code
591
559
 
592
560
  ``` r
593
561
  my_var <- x
594
562
  filter(df, my_var == 1)
595
563
  ```
596
- Generates the following error: "object 'x' not found.
597
564
 
598
- However, in Galaaz, arguments are referencially transparent as can be seen by the
599
- code below. Note initially that 'my_var = R[:x]' will not give the error "object 'x' not found"
600
- since ':x' is treated as an expression and assigned to my\_var. Then when doing (my\_var.eq 1),
601
- my\_var is a variable that resolves to ':x' and it becomes equivalent to (R[:x].eq 1) which is
602
- what we want.
565
+ Generates the following error: “object ‘x’ not found.
603
566
 
567
+ However, in Galaaz, arguments are referencially transparent as can be
568
+ seen by the code below. Note initially that ‘my_var = R\[:x\]’ will not
569
+ give the error “object ‘x’ not found” since ‘:x’ is treated as an
570
+ expression and assigned to my_var. Then when doing (my_var.eq 1), my_var
571
+ is a variable that resolves to ‘:x’ and it becomes equivalent to
572
+ (R\[:x\].eq 1) which is what we want.
604
573
 
605
574
  ``` ruby
606
575
  my_var = R[:x]
607
576
  puts df.filter(my_var.eq 1)
608
577
  ```
609
578
 
610
- ```
611
- ## x y
612
- ## 1 1 3
613
- ```
579
+ ## x y
580
+ ## 1 1 3
581
+
614
582
  As stated by Hadley
615
583
 
616
- > dplyr code is ambiguous. Depending on what variables are defined where,
617
- > filter(df, x == y) could be equivalent to any of:
584
+ > dplyr code is ambiguous. Depending on what variables are defined
585
+ > where, filter(df, x == y) could be equivalent to any of:
618
586
 
619
- ```
620
- df[df$x == df$y, ]
621
- df[df$x == y, ]
622
- df[x == df$y, ]
623
- df[x == y, ]
624
- ```
625
- In galaaz this ambiguity does not exist, filter(df, x.eq y) is not a valid expression as
626
- expressions are build with symbols. In doing filter(df, R[:x].eq y) we are looking for elements
627
- of the 'x' column that are equal to a previously defined y variable. Finally in
628
- filter(df, R[:x].eq R[:y]) we are looking for elements in which the 'x' column value is equal to
629
- the 'y' column value. This can be seen in the following two chunks of code:
587
+ df[df$x == df$y, ]
588
+ df[df$x == y, ]
589
+ df[x == df$y, ]
590
+ df[x == y, ]
630
591
 
592
+ In galaaz this ambiguity does not exist, filter(df, x.eq y) is not a
593
+ valid expression as expressions are build with symbols. In doing
594
+ filter(df, R\[:x\].eq y) we are looking for elements of the ‘x’ column
595
+ that are equal to a previously defined y variable. Finally in filter(df,
596
+ R\[:x\].eq R\[:y\]) we are looking for elements in which the ‘x’ column
597
+ value is equal to the ‘y’ column value. This can be seen in the
598
+ following two chunks of code:
631
599
 
632
600
  ``` ruby
633
601
  y = 1
@@ -637,11 +605,8 @@ x = 2
637
605
  puts df.filter(R[:x].eq R[:y])
638
606
  ```
639
607
 
640
- ```
641
- ## x y
642
- ## 1 2 2
643
- ```
644
-
608
+ ## x y
609
+ ## 1 2 2
645
610
 
646
611
  ``` ruby
647
612
  # looking for values where the 'x' column is equal to the 'y' variable
@@ -649,81 +614,82 @@ puts df.filter(R[:x].eq R[:y])
649
614
  puts df.filter(R[:x].eq y)
650
615
  ```
651
616
 
652
- ```
653
- ## x y
654
- ## 1 1 3
655
- ```
617
+ ## x y
618
+ ## 1 1 3
619
+
656
620
  ## Writing a function that applies to different data sets
657
621
 
658
- Let's suppose that we want to write a function that receives as the first argument a data frame
659
- and as second argument an expression that adds a column to the data frame that is equal to the
660
- sum of elements in column 'a' plus 'x'.
622
+ Let’s suppose that we want to write a function that receives as the
623
+ first argument a data frame and as second argument an expression that
624
+ adds a column to the data frame that is equal to the sum of elements in
625
+ column ‘a’ plus ‘x’.
661
626
 
662
- Here is the intended behaviour using the 'mutate' function of 'dplyr':
627
+ Here is the intended behaviour using the ‘mutate’ function of ‘dplyr’:
628
+
629
+ mutate(df1, y = a + x)
630
+ mutate(df2, y = a + x)
631
+ mutate(df3, y = a + x)
632
+ mutate(df4, y = a + x)
663
633
 
664
- ```
665
- mutate(df1, y = a + x)
666
- mutate(df2, y = a + x)
667
- mutate(df3, y = a + x)
668
- mutate(df4, y = a + x)
669
- ```
670
634
  The naive approach to writing an R function to solve this problem is:
671
635
 
672
- ```
673
- mutate_y <- function(df) {
674
- mutate(df, y = a + x)
675
- }
676
- ```
677
- Unfortunately, in R, this function can fail silently if one of the variables isn’t present
678
- in the data frame, but is present in the global environment. We will not go through here how
679
- to solve this problem in R.
636
+ mutate_y <- function(df) {
637
+ mutate(df, y = a + x)
638
+ }
680
639
 
681
- In Galaaz the method mutate_y below will work fine and will never fail silently.
640
+ Unfortunately, in R, this function can fail silently if one of the
641
+ variables isn’t present in the data frame, but is present in the global
642
+ environment. We will not go through here how to solve this problem in R.
682
643
 
644
+ In Galaaz the method mutate_y below will work fine and will never fail
645
+ silently.
683
646
 
684
647
  ``` ruby
685
648
  def mutate_y(df)
686
- # Mutate column names are Ruby kwargs (y: …). Use .assign only for R `<-` expressions.
649
+ # Column names are Ruby kwargs (y: …).
650
+ # Use .assign only for R `<-` expressions.
687
651
  df.mutate(y: R[:a] + R[:x])
688
652
  end
689
653
  ```
690
- Here we create a data frame that has only one column named 'x':
691
654
 
655
+ Here we create a data frame that has only one column named ‘x’:
692
656
 
693
657
  ``` ruby
694
658
  df1 = R.data__frame(x: (1..3))
695
659
  puts df1
696
660
  ```
697
661
 
698
- ```
699
- ## x
700
- ## 1 1
701
- ## 2 2
702
- ## 3 3
703
- ```
704
-
705
- Note that method mutate_y will fail independetly from the fact that variable 'a' is defined and
706
- in the scope of the method. Variable 'a' has no relationship with the symbol `R[:a]` used in the
707
- definition of 'mutate\_y' above:
662
+ ## x
663
+ ## 1 1
664
+ ## 2 2
665
+ ## 3 3
708
666
 
667
+ Note that method mutate_y will fail independetly from the fact that
668
+ variable ‘a’ is defined and in the scope of the method. Variable ‘a’ has
669
+ no relationship with the symbol `R[:a]` used in the definition of
670
+ ‘mutate_y’ above:
709
671
 
710
672
  ``` ruby
711
673
  a = 10
712
- mutate_y(df1)
674
+ begin
675
+ mutate_y(df1)
676
+ rescue => e
677
+ # Short message only — full backtraces overflow PDF code boxes.
678
+ puts "#{e.class}: #{e.message}"
679
+ end
713
680
  ```
714
681
 
715
- ```
716
- ## Error: ℹ In argument: `y = a + x`.
717
- ## Caused by error:
718
- ## ! object 'a' not found
719
- ```
720
- ## Different expressions
682
+ ## NewBridge::SessionClient::RProcessError: Error: ℹ In argument: `y = a + x`.
683
+ ## Caused by error:
684
+ ## ! object 'a' not found
721
685
 
722
- Let's move to the next problem as presented by Hadley where trying to write a function in R
723
- that will receive two argumens, the first a variable and the second an expression is not trivial.
724
- Below we create a data frame and we want to write a function that groups data by a variable and
725
- summarises it by an expression:
686
+ ## Different expressions
726
687
 
688
+ Let’s move to the next problem as presented by Hadley where trying to
689
+ write a function in R that will receive two argumens, the first a
690
+ variable and the second an expression is not trivial. Below we create a
691
+ data frame and we want to write a function that groups data by a
692
+ variable and summarises it by an expression:
727
693
 
728
694
  ``` r
729
695
  set.seed(123)
@@ -738,46 +704,39 @@ df <- data.frame(
738
704
  as.data.frame(df)
739
705
  ```
740
706
 
741
- ```
742
- ## g1 g2 a b
743
- ## 1 1 1 3 3
744
- ## 2 1 2 2 1
745
- ## 3 2 1 5 2
746
- ## 4 2 2 4 5
747
- ## 5 2 1 1 4
748
- ```
707
+ ## g1 g2 a b
708
+ ## 1 1 1 3 3
709
+ ## 2 1 2 2 1
710
+ ## 3 2 1 5 2
711
+ ## 4 2 2 4 5
712
+ ## 5 2 1 1 4
749
713
 
750
714
  ``` r
751
715
  d2 <- df %>%
752
716
  group_by(g1) %>%
753
717
  summarise(a = mean(a))
754
-
718
+
755
719
  as.data.frame(d2)
756
720
  ```
757
721
 
758
- ```
759
- ## g1 a
760
- ## 1 1 2.500000
761
- ## 2 2 3.333333
762
- ```
722
+ ## g1 a
723
+ ## 1 1 2.500000
724
+ ## 2 2 3.333333
763
725
 
764
726
  ``` r
765
727
  d2 <- df %>%
766
728
  group_by(g2) %>%
767
729
  summarise(a = mean(a))
768
-
769
- as.data.frame(d2)
730
+
731
+ as.data.frame(d2)
770
732
  ```
771
733
 
772
- ```
773
- ## g2 a
774
- ## 1 1 3
775
- ## 2 2 3
776
- ```
734
+ ## g2 a
735
+ ## 1 1 3
736
+ ## 2 2 3
777
737
 
778
738
  As shown by Hadley, one might expect this function to do the trick:
779
739
 
780
-
781
740
  ``` r
782
741
  my_summarise <- function(df, group_var) {
783
742
  df %>%
@@ -789,15 +748,17 @@ my_summarise <- function(df, group_var) {
789
748
  #> Error: Column `group_var` is unknown
790
749
  ```
791
750
 
792
- In order to solve this problem, coding with dplyr requires the introduction of many new concepts
793
- and functions such as 'quo', 'quos', 'enquo', 'enquos', '!!' (bang bang), '!!!' (triple bang).
794
- Again, we'll leave to Hadley the explanation on how to use all those functions.
795
-
796
- Now, let's try to implement the same function in galaaz. The next code block first prints the
797
- 'df' data frame defined previously in R (to access an R variable from Galaaz, we use the tilde
798
- operator '~' applied to the R variable name as symbol, i.e., ':df'. We then create the
799
- 'my_summarize' method and call it passing the R data frame and the group by variable ':g1':
751
+ In order to solve this problem, coding with dplyr requires the
752
+ introduction of many new concepts and functions such as ‘quo’, ‘quos’,
753
+ ‘enquo’, ‘enquos’, ‘!!’ (bang bang), ‘!!!’ (triple bang). Again, we’ll
754
+ leave to Hadley the explanation on how to use all those functions.
800
755
 
756
+ Now, let’s try to implement the same function in galaaz. The next code
757
+ block first prints the ‘df’ data frame defined previously in R (to
758
+ access an R variable from Galaaz, we use the tilde operator ‘~’ applied
759
+ to the R variable name as symbol, i.e., ‘:df’. We then create the
760
+ ‘my_summarize’ method and call it passing the R data frame and the group
761
+ by variable ‘:g1’:
801
762
 
802
763
  ``` ruby
803
764
  puts ~R[:df]
@@ -812,64 +773,59 @@ end
812
773
  puts my_summarize(~R[:df], R[:g1])
813
774
  ```
814
775
 
815
- ```
816
- ## g1 g2 a b
817
- ## 1 1 1 3 3
818
- ## 2 1 2 2 1
819
- ## 3 2 1 5 2
820
- ## 4 2 2 4 5
821
- ## 5 2 1 1 4
822
- ##
823
- ## # A tibble: 2 × 2
824
- ## g1 a
825
- ## <dbl> <dbl>
826
- ## 1 1 2.5
827
- ## 2 2 3.33
828
- ```
829
- It works!!! Well, let's make sure this was not just some coincidence
776
+ ## g1 g2 a b
777
+ ## 1 1 1 3 3
778
+ ## 2 1 2 2 1
779
+ ## 3 2 1 5 2
780
+ ## 4 2 2 4 5
781
+ ## 5 2 1 1 4
782
+ ##
783
+ ## # A tibble: 2 × 2
784
+ ## g1 a
785
+ ## <dbl> <dbl>
786
+ ## 1 1 2.5
787
+ ## 2 2 3.33
830
788
 
789
+ It works!!! Well, let’s make sure this was not just some coincidence
831
790
 
832
791
  ``` ruby
833
792
  puts my_summarize(~R[:df], R[:g2])
834
793
  ```
835
794
 
836
- ```
837
- ## # A tibble: 2 × 2
838
- ## g2 a
839
- ## <dbl> <dbl>
840
- ## 1 1 3
841
- ## 2 2 3
842
- ```
795
+ ## # A tibble: 2 × 2
796
+ ## g2 a
797
+ ## <dbl> <dbl>
798
+ ## 1 1 3
799
+ ## 2 2 3
843
800
 
844
- Great, everything is fine! No magic, no new functions, no complexities, just normal, standard Ruby
845
- code. If you've ever done NSE in R, this certainly feels much safer and easy to implement.
801
+ Great, everything is fine! No magic, no new functions, no complexities,
802
+ just normal, standard Ruby code. If you’ve ever done NSE in R, this
803
+ certainly feels much safer and easy to implement.
846
804
 
847
805
  ## Different input variables
848
806
 
849
- In the previous section we've managed to get rid of all NSE formulation for a simple example, but
850
- does this remain true for more complex examples, or will the Galaaz way prove inpractical for
851
- more complex code?
807
+ In the previous section we’ve managed to get rid of all NSE formulation
808
+ for a simple example, but does this remain true for more complex
809
+ examples, or will the Galaaz way prove inpractical for more complex
810
+ code?
852
811
 
853
- In the next example Hadley proposes us to write a function that given an expression such as 'a'
854
- or 'a * b', calculates three summaries. What we want a function that does the same as these R
855
- statements:
812
+ In the next example Hadley proposes us to write a function that given an
813
+ expression such as ‘a’ or ‘a \* b’, calculates three summaries. What we
814
+ want a function that does the same as these R statements:
856
815
 
857
- ```
858
- summarise(df, mean = mean(a), sum = sum(a), n = n())
859
- #> # A tibble: 1 x 3
860
- #> mean sum n
861
- #> <dbl> <int> <int>
862
- #> 1 3 15 5
863
-
864
- summarise(df, mean = mean(a * b), sum = sum(a * b), n = n())
865
- #> # A tibble: 1 x 3
866
- #> mean sum n
867
- #> <dbl> <int> <int>
868
- #> 1 9 45 5
869
- ```
816
+ summarise(df, mean = mean(a), sum = sum(a), n = n())
817
+ #> # A tibble: 1 x 3
818
+ #> mean sum n
819
+ #> <dbl> <int> <int>
820
+ #> 1 3 15 5
870
821
 
871
- Let's try it in galaaz:
822
+ summarise(df, mean = mean(a * b), sum = sum(a * b), n = n())
823
+ #> # A tibble: 1 x 3
824
+ #> mean sum n
825
+ #> <dbl> <int> <int>
826
+ #> 1 9 45 5
872
827
 
828
+ Let’s try it in galaaz:
873
829
 
874
830
  ``` ruby
875
831
  def my_summarise2(df, expr)
@@ -884,50 +840,49 @@ puts my_summarise2((~R[:df]), :a)
884
840
  puts my_summarise2((~R[:df]), R[:a] * R[:b])
885
841
  ```
886
842
 
887
- ```
888
- ## mean sum n
889
- ## 1 3 15 5
890
- ## mean sum n
891
- ## 1 9 45 5
892
- ```
843
+ ## mean sum n
844
+ ## 1 3 15 5
845
+ ## mean sum n
846
+ ## 1 9 45 5
893
847
 
894
- Once again, there is no need to use any special theory or functions. The only point to be
895
- careful about is the use of 'E' to build expressions from functions 'mean', 'sum' and 'n'.
848
+ Once again, there is no need to use any special theory or functions. The
849
+ only point to be careful about is the use of ‘E’ to build expressions
850
+ from functions ‘mean’, ‘sum’ and ‘n’.
896
851
 
897
852
  ## Different input and output variable
898
853
 
899
- Now the next challenge presented by Hadley is to vary the name of the output variables based on
900
- the received expression. So, if the input expression is 'a', we want our data frame columns to
901
- be named 'mean\_a' and 'sum\_a'. Now, if the input expression is 'b', columns
902
- should be named 'mean\_b' and 'sum\_b'.
903
-
904
- ```
905
- mutate(df, mean_a = mean(a), sum_a = sum(a))
906
- #> # A tibble: 5 x 6
907
- #> g1 g2 a b mean_a sum_a
908
- #> <dbl> <dbl> <int> <int> <dbl> <int>
909
- #> 1 1 1 1 3 3 15
910
- #> 2 1 2 4 2 3 15
911
- #> 3 2 1 2 1 3 15
912
- #> 4 2 2 5 4 3 15
913
- #> # … with 1 more row
914
-
915
- mutate(df, mean_b = mean(b), sum_b = sum(b))
916
- #> # A tibble: 5 x 6
917
- #> g1 g2 a b mean_b sum_b
918
- #> <dbl> <dbl> <int> <int> <dbl> <int>
919
- #> 1 1 1 1 3 3 15
920
- #> 2 1 2 4 2 3 15
921
- #> 3 2 1 2 1 3 15
922
- #> 4 2 2 5 4 3 15
923
- #> # … with 1 more row
924
- ```
925
- In order to solve this problem in R, Hadley needs to introduce some more new functions and notations:
926
- 'quo_name' and the ':=' operator from package 'rlang'
854
+ Now the next challenge presented by Hadley is to vary the name of the
855
+ output variables based on the received expression. So, if the input
856
+ expression is ‘a’, we want our data frame columns to be named ‘mean_a’
857
+ and ‘sum_a’. Now, if the input expression is ‘b’, columns should be
858
+ named ‘mean_b’ and ‘sum_b’.
859
+
860
+ mutate(df, mean_a = mean(a), sum_a = sum(a))
861
+ #> # A tibble: 5 x 6
862
+ #> g1 g2 a b mean_a sum_a
863
+ #> <dbl> <dbl> <int> <int> <dbl> <int>
864
+ #> 1 1 1 1 3 3 15
865
+ #> 2 1 2 4 2 3 15
866
+ #> 3 2 1 2 1 3 15
867
+ #> 4 2 2 5 4 3 15
868
+ #> # … with 1 more row
869
+
870
+ mutate(df, mean_b = mean(b), sum_b = sum(b))
871
+ #> # A tibble: 5 x 6
872
+ #> g1 g2 a b mean_b sum_b
873
+ #> <dbl> <dbl> <int> <int> <dbl> <int>
874
+ #> 1 1 1 1 3 3 15
875
+ #> 2 1 2 4 2 3 15
876
+ #> 3 2 1 2 1 3 15
877
+ #> 4 2 2 5 4 3 15
878
+ #> # … with 1 more row
879
+
880
+ In order to solve this problem in R, Hadley needs to introduce some more
881
+ new functions and notations: ‘quo_name’ and the ‘:=’ operator from
882
+ package ‘rlang’
927
883
 
928
884
  Here is our Ruby code:
929
885
 
930
-
931
886
  ``` ruby
932
887
  def my_mutate(df, expr)
933
888
  mean_name = "mean_#{expr.to_s}"
@@ -941,36 +896,37 @@ puts my_mutate((~R[:df]), :a)
941
896
  puts my_mutate((~R[:df]), :b)
942
897
  ```
943
898
 
944
- ```
945
- ## g1 g2 a b mean_a sum_a
946
- ## 1 1 1 3 3 3 15
947
- ## 2 1 2 2 1 3 15
948
- ## 3 2 1 5 2 3 15
949
- ## 4 2 2 4 5 3 15
950
- ## 5 2 1 1 4 3 15
951
- ## g1 g2 a b mean_b sum_b
952
- ## 1 1 1 3 3 3 15
953
- ## 2 1 2 2 1 3 15
954
- ## 3 2 1 5 2 3 15
955
- ## 4 2 2 4 5 3 15
956
- ## 5 2 1 1 4 3 15
957
- ```
958
- It really seems that "Non Standard Evaluation" is actually quite standard in Galaaz! But, you
959
- might have noticed a small change in the way the arguments to the mutate method were called.
960
- In a previous example we used df.summarise(mean: E.mean(:a), ...) where the column name was
961
- followed by a ':' colom. In this example, we have df.mutate(mean_name => E.mean(expr), ...)
962
- and variable mean\_name is not followed by ':' but by '=>'. This is standard Ruby notation.
963
-
964
- [explain....]
899
+ ## g1 g2 a b mean_a sum_a
900
+ ## 1 1 1 3 3 3 15
901
+ ## 2 1 2 2 1 3 15
902
+ ## 3 2 1 5 2 3 15
903
+ ## 4 2 2 4 5 3 15
904
+ ## 5 2 1 1 4 3 15
905
+ ## g1 g2 a b mean_b sum_b
906
+ ## 1 1 1 3 3 3 15
907
+ ## 2 1 2 2 1 3 15
908
+ ## 3 2 1 5 2 3 15
909
+ ## 4 2 2 4 5 3 15
910
+ ## 5 2 1 1 4 3 15
911
+
912
+ It really seems that “Non Standard Evaluation” is actually quite
913
+ standard in Galaaz! But, you might have noticed a small change in the
914
+ way the arguments to the mutate method were called. In a previous
915
+ example we used df.summarise(mean: E.mean(:a), …) where the column name
916
+ was followed by a ‘:’ colom. In this example, we have
917
+ df.mutate(mean_name =\> E.mean(expr), …) and variable mean_name is not
918
+ followed by ‘:’ but by ‘=\>’. This is standard Ruby notation.
919
+
920
+ \[explain….\]
965
921
 
966
922
  ## Capturing multiple variables
967
923
 
968
- Moving on with new complexities, Hadley proposes us to solve the problem in which the
969
- summarise function will receive any number of grouping variables.
970
-
971
- This again is quite standard Ruby. In order to receive an undefined number of paramenters
972
- the paramenter is preceded by '*':
924
+ Moving on with new complexities, Hadley proposes us to solve the problem
925
+ in which the summarise function will receive any number of grouping
926
+ variables.
973
927
 
928
+ This again is quite standard Ruby. In order to receive an undefined
929
+ number of paramenters the paramenter is preceded by ’\*’:
974
930
 
975
931
  ``` ruby
976
932
  def my_summarise3(df, *group_vars)
@@ -981,85 +937,89 @@ end
981
937
  puts my_summarise3((~R[:df]), :g1, :g2)
982
938
  ```
983
939
 
984
- ```
985
- ## # A tibble: 4 × 3
986
- ## # Groups: g1 [2]
987
- ## g1 g2 a
988
- ## <dbl> <dbl> <dbl>
989
- ## 1 1 1 3
990
- ## 2 1 2 2
991
- ## 3 2 1 3
992
- ## 4 2 2 4
993
- ```
940
+ ## # A tibble: 4 × 3
941
+ ## # Groups: g1 [2]
942
+ ## g1 g2 a
943
+ ## <dbl> <dbl> <dbl>
944
+ ## 1 1 1 3
945
+ ## 2 1 2 2
946
+ ## 3 2 1 3
947
+ ## 4 2 2 4
994
948
 
995
949
  # Why does R require NSE and Galaaz does not?
996
950
 
997
- NSE introduces a number of new concepts, such as 'quoting', 'quasiquotation', 'unquoting' and
998
- 'unquote-splicing', while in Galaaz none of those concepts are needed. What gives?
999
-
1000
- R is an extremely flexible language and it has lazy evaluation of parameters. When in R a
1001
- function is called as 'summarise(df, a = b)', the summarise function receives the litteral
1002
- 'a = b' parameter and can work with this as if it were a string. In R, it is not clear what
1003
- a and b are, they can be expressions or they can be variables, it is up to the function to
1004
- decide what 'a = b' means.
1005
-
1006
- In Ruby, there is no lazy evaluation of parameters and 'a' is always a variable and so is 'b'.
1007
- Variables assume their value as soon as they are used, so 'x = a' is immediately evaluate and
1008
- variable 'x' will receive the value of variable 'a' as soon as the Ruby statement is executed.
1009
- Ruby also provides the notion of a symbol; ':a' is a symbol and does not evaluate to anything.
1010
- Galaaz uses Ruby symbols to build expressions that are not bound to anything: 'R[:a].eq R[:b]' is
1011
- clearly an expression and has no relationship whatsoever with the statment 'a = b'. By using
1012
- symbols, variables and expressions all the possible ambiguities that are found in R are
1013
- eliminated in Galaaz.
1014
-
1015
- The main problem that remains, is that in R, functions are not clearly documented as what type
1016
- of input they are expecting, they might be expecting regular variables or they might be
1017
- expecting expressions and the R function will know how to deal with an input of the form
1018
- 'a = b', now for the Ruby developer it might not be immediately clear if it should call the
1019
- function passing the value 'true' if variable 'a' is equal to variable 'b' or if it should
1020
- call the function passing the expression 'R[:a].eq R[:b]'.
1021
-
951
+ NSE introduces a number of new concepts, such as ‘quoting’,
952
+ ‘quasiquotation’, ‘unquoting’ and ‘unquote-splicing’, while in Galaaz
953
+ none of those concepts are needed. What gives?
954
+
955
+ R is an extremely flexible language and it has lazy evaluation of
956
+ parameters. When in R a function is called as ‘summarise(df, a = b)’,
957
+ the summarise function receives the litteral ‘a = b’ parameter and can
958
+ work with this as if it were a string. In R, it is not clear what a and
959
+ b are, they can be expressions or they can be variables, it is up to the
960
+ function to decide what ‘a = b’ means.
961
+
962
+ In Ruby, there is no lazy evaluation of parameters and ‘a’ is always a
963
+ variable and so is ‘b’. Variables assume their value as soon as they are
964
+ used, so ‘x = a’ is immediately evaluate and variable ‘x’ will receive
965
+ the value of variable ‘a’ as soon as the Ruby statement is executed.
966
+ Ruby also provides the notion of a symbol; ‘:a’ is a symbol and does not
967
+ evaluate to anything. Galaaz uses Ruby symbols to build expressions that
968
+ are not bound to anything: ‘R\[:a\].eq R\[:b\]’ is clearly an expression
969
+ and has no relationship whatsoever with the statment ‘a = b’. By using
970
+ symbols, variables and expressions all the possible ambiguities that are
971
+ found in R are eliminated in Galaaz.
972
+
973
+ The main problem that remains, is that in R, functions are not clearly
974
+ documented as what type of input they are expecting, they might be
975
+ expecting regular variables or they might be expecting expressions and
976
+ the R function will know how to deal with an input of the form ‘a = b’,
977
+ now for the Ruby developer it might not be immediately clear if it
978
+ should call the function passing the value ‘true’ if variable ‘a’ is
979
+ equal to variable ‘b’ or if it should call the function passing the
980
+ expression ‘R\[:a\].eq R\[:b\]’.
1022
981
 
1023
982
  # Advanced dplyr features
1024
983
 
1025
- In the blog: [Programming with dplyr by using dplyr](https://www.r-bloggers.com/programming-with-dplyr-by-using-dplyr/) Iñaki Úcar shows surprise that some R users are trying to code in dplyr avoiding
1026
- the use of NSE. For instance he says:
1027
-
1028
- > Take the example of seplyr. It stands for standard evaluation dplyr, and enables us to
1029
- > program over dplyr without having “to bring in (or study) any deep-theory or
1030
- > heavy-weight tools such as rlang/tidyeval”.
984
+ In the blog: [Programming with dplyr by using
985
+ dplyr](https://www.r-bloggers.com/programming-with-dplyr-by-using-dplyr/)
986
+ Iñaki Úcar shows surprise that some R users are trying to code in dplyr
987
+ avoiding the use of NSE. For instance he says:
1031
988
 
1032
- For me, there isn't really any surprise that users are trying to avoid dplyr deep-theory. R
1033
- users frequently are not programmers and learning to code is already hard business, on top
1034
- of that, having to learn how to 'quote' or 'enquo' or 'quos' or 'enquos' is not necessarily
1035
- a 'piece of cake'. So much so, that 'tidyeval' has some more advanced functions that instead
1036
- of using quoted expressions, uses strings as arguments.
989
+ > Take the example of seplyr. It stands for standard evaluation dplyr,
990
+ > and enables us to program over dplyr without having “to bring in (or
991
+ > study) any deep-theory or heavy-weight tools such as rlang/tidyeval”.
1037
992
 
1038
- In the following examples, we show the use of functions 'group\_by\_at', 'summarise\_at' and
1039
- 'rename\_at' that receive strings as argument. The data frame used in 'starwars' that describes
1040
- features of characters in the Starwars movies:
993
+ For me, there isn’t really any surprise that users are trying to avoid
994
+ dplyr deep-theory. R users frequently are not programmers and learning
995
+ to code is already hard business, on top of that, having to learn how to
996
+ ‘quote’ or ‘enquo’ or ‘quos’ or ‘enquos’ is not necessarily a ‘piece of
997
+ cake’. So much so, that ‘tidyeval’ has some more advanced functions that
998
+ instead of using quoted expressions, uses strings as arguments.
1041
999
 
1000
+ In the following examples, we show the use of functions ‘group_by_at’,
1001
+ ‘summarise_at’ and ‘rename_at’ that receive strings as argument. The
1002
+ data frame used in ‘starwars’ that describes features of characters in
1003
+ the Starwars movies:
1042
1004
 
1043
1005
  ``` ruby
1044
1006
  puts (~R[:starwars]).head
1045
1007
  ```
1046
1008
 
1047
- ```
1048
- ## # A tibble: 6 × 14
1049
- ## name height mass hair_color skin_color eye_color birth_year sex gender
1050
- ## <chr> <int> <dbl> <chr> <chr> <chr> <dbl> <chr> <chr>
1051
- ## 1 Luke Sky… 172 77 blond fair blue 19 male mascu…
1052
- ## 2 C-3PO 167 75 <NA> gold yellow 112 none mascu…
1053
- ## 3 R2-D2 96 32 <NA> white, bl… red 33 none mascu…
1054
- ## 4 Darth Va… 202 136 none white yellow 41.9 male mascu…
1055
- ## 5 Leia Org… 150 49 brown light brown 19 fema… femin…
1056
- ## 6 Owen Lars 178 120 brown, gr… light blue 52 male mascu…
1057
- ## # ℹ 5 more variables: homeworld <chr>, species <chr>, films <list>,
1058
- ## # vehicles <list>, starships <list>
1059
- ```
1060
- The grouped_mean function below will receive a grouping variable and calculate summaries for
1061
- the value\_variables given:
1009
+ ## # A tibble: 6 × 14
1010
+ ## name height mass hair_color skin_color eye_color birth_year sex
1011
+ ## <chr> <int> <dbl> <chr> <chr> <chr> <dbl> <chr>
1012
+ ## 1 Luke … 172 77 blond fair blue 19 male
1013
+ ## 2 C-3PO 167 75 <NA> gold yellow 112 none
1014
+ ## 3 R2-D2 96 32 <NA> white, bl… red 33 none
1015
+ ## 4 Darth… 202 136 none white yellow 41.9 male
1016
+ ## 5 Leia … 150 49 brown light brown 19 fema…
1017
+ ## 6 Owen … 178 120 brown, gr… light blue 52 male
1018
+ ## # ℹ 6 more variables: gender <chr>, homeworld <chr>, species <chr>,
1019
+ ## # films <list>, vehicles <list>, starships <list>
1062
1020
 
1021
+ The grouped_mean function below will receive a grouping variable and
1022
+ calculate summaries for the value_variables given:
1063
1023
 
1064
1024
  ``` r
1065
1025
  grouped_mean <- function(data, grouping_variables, value_variables) {
@@ -1074,94 +1034,105 @@ gm = starwars %>%
1074
1034
  grouped_mean("eye_color", c("mass", "birth_year"))
1075
1035
  ```
1076
1036
 
1077
- ```
1078
- ## Warning: `funs()` was deprecated in dplyr 0.8.0.
1079
- ## ℹ Please use a list of either functions or lambdas:
1080
- ##
1081
- ## # Simple named list: list(mean = mean, median = median)
1082
- ##
1083
- ## # Auto named with `tibble::lst()`: tibble::lst(mean, median)
1084
- ##
1085
- ## # Using lambdas list(~ mean(., trim = .2), ~ median(., na.rm = TRUE))
1086
- ## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
1087
- ## generated.
1088
- ```
1037
+ ## Warning: `funs()` was deprecated in dplyr 0.8.0.
1038
+ ## ℹ Please use a list of either functions or lambdas:
1039
+ ##
1040
+ ## # Simple named list: list(mean = mean, median = median)
1041
+ ##
1042
+ ## # Auto named with `tibble::lst()`: tibble::lst(mean, median)
1043
+ ##
1044
+ ## # Using lambdas list(~ mean(., trim = .2), ~ median(., na.rm = TRUE))
1045
+ ## Call `lifecycle::last_lifecycle_warnings()` to see where this warning
1046
+ ## was generated.
1089
1047
 
1090
1048
  ``` r
1091
1049
  as.data.frame(gm)
1092
1050
  ```
1093
1051
 
1094
- ```
1095
- ## eye_color mean_mass mean_birth_year count
1096
- ## 1 black 76.28571 33.00000 10
1097
- ## 2 blue 86.51667 67.06923 19
1098
- ## 3 blue-gray 77.00000 57.00000 1
1099
- ## 4 brown 66.09231 108.96429 21
1100
- ## 5 dark NaN NaN 1
1101
- ## 6 gold NaN NaN 1
1102
- ## 7 green, yellow 159.00000 NaN 1
1103
- ## 8 hazel 66.00000 34.50000 3
1104
- ## 9 orange 282.33333 231.00000 8
1105
- ## 10 pink NaN NaN 1
1106
- ## 11 red 81.40000 33.66667 5
1107
- ## 12 red, blue NaN NaN 1
1108
- ## 13 unknown 31.50000 NaN 3
1109
- ## 14 white 48.00000 NaN 1
1110
- ## 15 yellow 81.11111 76.38000 11
1111
- ```
1052
+ ## eye_color mean_mass mean_birth_year count
1053
+ ## 1 black 76.28571 33.00000 10
1054
+ ## 2 blue 86.51667 67.06923 19
1055
+ ## 3 blue-gray 77.00000 57.00000 1
1056
+ ## 4 brown 66.09231 108.96429 21
1057
+ ## 5 dark NaN NaN 1
1058
+ ## 6 gold NaN NaN 1
1059
+ ## 7 green, yellow 159.00000 NaN 1
1060
+ ## 8 hazel 66.00000 34.50000 3
1061
+ ## 9 orange 282.33333 231.00000 8
1062
+ ## 10 pink NaN NaN 1
1063
+ ## 11 red 81.40000 33.66667 5
1064
+ ## 12 red, blue NaN NaN 1
1065
+ ## 13 unknown 31.50000 NaN 3
1066
+ ## 14 white 48.00000 NaN 1
1067
+ ## 15 yellow 81.11111 76.38000 11
1112
1068
 
1113
1069
  The same code with Galaaz, becomes:
1114
1070
 
1115
-
1116
1071
  ``` ruby
1117
1072
  def grouped_mean(data, grouping_variables, value_variables)
1118
1073
  data.
1119
1074
  group_by_at(grouping_variables).
1120
1075
  mutate(count: E.n).
1121
- summarise_at(E.c(value_variables, "count"), R[:mean], na__rm: true).
1122
- rename_at(value_variables, E.funs(E.paste0("mean_", value_variables)))
1076
+ summarise_at(
1077
+ E.c(value_variables, "count"),
1078
+ R[:mean],
1079
+ na__rm: true).
1080
+ rename_at(
1081
+ value_variables,
1082
+ E.funs(E.paste0("mean_", value_variables)))
1123
1083
  end
1124
1084
 
1125
- puts grouped_mean((~R[:starwars]), "eye_color", E.c("mass", "birth_year"))
1126
- ```
1127
-
1128
- ```
1129
- ## # A tibble: 15 × 4
1130
- ## eye_color mean_mass mean_birth_year count
1131
- ## <chr> <dbl> <dbl> <dbl>
1132
- ## 1 black 76.3 33 10
1133
- ## 2 blue 86.5 67.1 19
1134
- ## 3 blue-gray 77 57 1
1135
- ## 4 brown 66.1 109. 21
1136
- ## 5 dark NaN NaN 1
1137
- ## 6 gold NaN NaN 1
1138
- ## 7 green, yellow 159 NaN 1
1139
- ## 8 hazel 66 34.5 3
1140
- ## 9 orange 282. 231 8
1141
- ## 10 pink NaN NaN 1
1142
- ## 11 red 81.4 33.7 5
1143
- ## 12 red, blue NaN NaN 1
1144
- ## 13 unknown 31.5 NaN 3
1145
- ## 14 white 48 NaN 1
1146
- ## 15 yellow 81.1 76.4 11
1147
- ```
1085
+ puts grouped_mean(
1086
+ (~R[:starwars]),
1087
+ "eye_color",
1088
+ E.c("mass", "birth_year"))
1089
+ ```
1090
+
1091
+ ## # A tibble: 15 × 4
1092
+ ## eye_color mean_mass mean_birth_year count
1093
+ ## <chr> <dbl> <dbl> <dbl>
1094
+ ## 1 black 76.3 33 10
1095
+ ## 2 blue 86.5 67.1 19
1096
+ ## 3 blue-gray 77 57 1
1097
+ ## 4 brown 66.1 109. 21
1098
+ ## 5 dark NaN NaN 1
1099
+ ## 6 gold NaN NaN 1
1100
+ ## 7 green, yellow 159 NaN 1
1101
+ ## 8 hazel 66 34.5 3
1102
+ ## 9 orange 282. 231 8
1103
+ ## 10 pink NaN NaN 1
1104
+ ## 11 red 81.4 33.7 5
1105
+ ## 12 red, blue NaN NaN 1
1106
+ ## 13 unknown 31.5 NaN 3
1107
+ ## 14 white 48 NaN 1
1108
+ ## 15 yellow 81.1 76.4 11
1148
1109
 
1149
1110
  # Further reading
1150
1111
 
1151
- * [JRuby](https://www.jruby.org/) — Ruby on the JVM (Galaaz 2.0)
1152
- * [How to make Beautiful Ruby Plots with Galaaz](https://medium.freecodecamp.org/how-to-make-beautiful-ruby-plots-with-galaaz-320848058857) (plots; narrative partly pre-2.0)
1153
- * [Ruby Plotting with Galaaz in GraalVM](https://towardsdatascience.com/ruby-plotting-with-galaaz-an-example-of-tightly-coupling-ruby-and-r-in-graalvm-520b69e21021) (older stack; ideas still useful)
1154
- * [How to do reproducible research in Ruby with gKnit](https://towardsdatascience.com/how-to-do-reproducible-research-in-ruby-with-gknit-c26d2684d64e)
1155
- * [R for Data Science](https://r4ds.had.co.nz/)
1156
- * [Advanced R](https://adv-r.hadley.nz/)
1157
- * Historical context: [GraalVM](https://www.graalvm.org/), [TruffleRuby](https://github.com/oracle/truffleruby), [FastR](https://github.com/oracle/fastr)
1112
+ - [JRuby](https://www.jruby.org/) — Ruby on the JVM (Galaaz 2.0)
1113
+ - [How to make Beautiful Ruby Plots with
1114
+ Galaaz](https://medium.freecodecamp.org/how-to-make-beautiful-ruby-plots-with-galaaz-320848058857)
1115
+ (plots; narrative partly pre-2.0)
1116
+ - [Ruby Plotting with Galaaz in
1117
+ GraalVM](https://towardsdatascience.com/ruby-plotting-with-galaaz-an-example-of-tightly-coupling-ruby-and-r-in-graalvm-520b69e21021)
1118
+ (older stack; ideas still useful)
1119
+ - [How to do reproducible research in Ruby with
1120
+ gKnit](https://towardsdatascience.com/how-to-do-reproducible-research-in-ruby-with-gknit-c26d2684d64e)
1121
+ - [R for Data Science](https://r4ds.had.co.nz/)
1122
+ - [Advanced R](https://adv-r.hadley.nz/)
1123
+ - Historical context: [GraalVM](https://www.graalvm.org/),
1124
+ [TruffleRuby](https://github.com/oracle/truffleruby),
1125
+ [FastR](https://github.com/oracle/fastr)
1158
1126
 
1159
1127
  # Conclusion
1160
1128
 
1161
- Ruby and Galaaz provide a nice framework for developing code that uses R functions. Although R is
1162
- a very powerful and flexible language, sometimes, too much flexibility makes life harder for
1163
- the casual user. We believe however, that even for the advanced user, Ruby integrated
1164
- with R throught Galaaz, makes a powerful environment for data analysis. In this blog post we
1165
- showed how Galaaz consistent syntax eliminates the need for complex constructs such as quoting,
1166
- enquoting, quasiquotation, etc. This simplification comes from the fact that expressions and
1167
- variables are clearly separated objects, which is not the case in the R language.
1129
+ Ruby and Galaaz provide a nice framework for developing code that uses R
1130
+ functions. Although R is a very powerful and flexible language,
1131
+ sometimes, too much flexibility makes life harder for the casual user.
1132
+ We believe however, that even for the advanced user, Ruby integrated
1133
+ with R throught Galaaz, makes a powerful environment for data analysis.
1134
+ In this blog post we showed how Galaaz consistent syntax eliminates the
1135
+ need for complex constructs such as quoting, enquoting, quasiquotation,
1136
+ etc. This simplification comes from the fact that expressions and
1137
+ variables are clearly separated objects, which is not the case in the R
1138
+ language.