Alex Rivera | Logout

Any hard data on GC vs explicit memory management performance?

Asked 2009-04-16T12:22:33.160
32

I recently read the excellent article "The Transactional Memory / Garbage Collection Analogy" by Dan Grossman. One sentence really caught my attention:

In theory, garbage collection can improve performance by increasing spatial locality (due to object-relocation), but in practice we pay a moderate performance cost for software engineering benefits.

Until then, my feeling had always been very vague about it. Over and over, you see claims that GC can be more efficient, so I always kept that notion in the back of my head. After reading this, however, I started having serious doubts.

As an experiment to measure the impact on GC languages, some people took some Java programs, traced the execution, and then replaced garbage collection with explicit memory management. According to this review of the article on Lambda the ultimate, they found out that GC was always slower. Virtual memory issues made GC look even worse, since the collector regularly touches way more memory pages than the program itself at that point, and therefore causes a lot of swapping.

This is all experimental to me. Has anybody, and in particular in the context of C++, performed a comprehensive benchmark of GC performance when comparing to explicit memory management?

Particularly interesting would be to compare how various big open-source projects, for example, perform with or without GC. Has anybody heard of such results before?

EDIT: And please focus on the performance problem, not on why GC exists or why it is beneficial.

Cheers,

Carl

PS. In case you're already pulling out the flame-thrower: I am not trying to disqualify GC, I'm just t

Edit
Report

1 Answer

-2

As @dribeas points out, the biggest 'confound' to the experiment in the (Hertz&Berger) paper is that code is always written under some 'implicit assumptions' about what is cheap and what is expensive. Apart from that confound, the experimental methodology (run a Java program offline, create an oracle of object lifetimes, instrument back in the 'ideal' alloc/free calls) is actually quite brilliant and illuminating . (And my personal opinion is that confound does not detract too much from their results.)

Personally, my gut-feel is that using a GC-ed runtime means accepting a factor-of-three performance hit to your application (GC'd will be 3x slower). But the real landscape of programs is littered with confounds, and you'd be likely to find a huge scatterplot of data if you could perform the 'ideal' experiment on lots of programs across many application domains, with GC sometimes winning and Manual often winning. (And the landscape is continually changing - will the results change when multicore (and software designed for multicore) is mainstream?)

See also my answer to

Are there statistical studies that indicates that Python is "more productive"?

which has the thesis that "due to so many confounds, all evidence about software engineering is anecdotal".

answered 2009-04-19T01:48:14.707

Your Answer