Alex Rivera | Logout

writing functions vs. line-by-line interpretation in an R workflow

Asked 2011-03-20T00:43:18.923
10

Much has been written here about developing a workflow in R for statistical projects. The most popular workflow seems to be Josh Reich's LCFD model. With a main.R containing code:

source('load.R')
source('clean.R')
source('func.R')
source('do.R')

so that a single source('main.R') runs the entire project.

Q: Is there a reason to prefer this workflow to one in which the line-by-line interpretive work done in load.R, clean.R, and do.R is replaced by functions which are called by main.R?

I can't find the link now, but I had read somewhere on SO that when programming in R one must get over their desire to write everything in terms of function calls---that R was MEANT to be written is this line-by-line interpretive form.

Q: Really? Why?

I've been frustrated with the LCFD approach and am going to probably write everything in terms of function calls. But before doing this, I'd like to hear from the good folks of SO as to whether this is a good idea or not.

EDIT: The project I'm working on right now is to (1) read in a set of financial data, (2) clean it (quite involved), (3) Estimate some quantity associated with the data using my estimator (4) Estimate that same quantity using traditional estimators (5) Report results. My programs should be written in such a way that it's a cinch to do the work (1) for different empirical data sets, (2) for simulation data, or (3) using different estimators. ALSO, it should follow literate programming and reproducible research guidelines so that it's simple for a newcomer to the code to run the program, understand what's going on, and how to tweak it.

Edit
Report

2 Answers

8

No one has mentioned an important consideration when writing functions: there's not much point in writing them unless you're repeating some action again and again. In some parts of an analysis, you'll being doing one-off operations, so there's not much point in writing a function for them. If you have to repeat something more than a few times, it's worth investing the time and effort to write a re-usable function.

answered 2011-03-21T03:58:45.610
6

Workflow:

I use something very similar:

  1. Base.r: pulls primary data, calls on other files (items 2 through 5)
  2. Functions.r: loads functions
  3. Plot Options.r: loads a number of general plot options I use frequently
  4. Lists.r: loads lists, I have a lot of them because company names, statements and the like change over time
  5. Recodes.r: most of the work is done in this file, essentially it's data cleaning and sorting

No analysis has been done up to this point. This is just for data cleaning and sorting.

At the end of Recodes.r I save the environment to be reloaded into my actual analysis.

save(list=ls(), file="Cleaned.Rdata")

With the cleaning done, functions and plot options ready, I start getting into my analysis. Again, I continue to break it up into smaller files that are focused into topics or themes, like: demographics, client requests, correlations, correspondence analysis, plots, ect. I almost always run the first 5 automatically to get my environment set up and then I run the others on a line by line basis to ensure accuracy and explore.

At the beginning of every file I load the cleaned data environment and prosper.

load("Cleaned.Rdata")

Object Nomenclature:

I don't use lists, but I do use a nomenclature for my objects.

df.YYYY # Data for a certain year
demo.describe.YYYY ## Demographic data for a certain year
po.describe ## Plot option
list.describe.YYYY ## lists
f.describe ## Functions

Using a friendly mnemonic to replace "describe" in the above.

Commenting

I've been trying to get myself into the habit of using comment(x) which I've found incredibly useful. Comments in the code are helpful but oftentimes not enough.

Cleaning Up

Again, here, I always try to use the same object(s) for easy cleanup. tmp, tmp1, tmp2, tmp3 for

answered 2011-03-22T01:35:07.510

Your Answer