6 ms·
Tune Code Before Your Garbage Collector
- peter_lawrey 2mo agoIn this post I look at a simple event to response latency benchmark, MarketDataSnapshot to NewOrderSingle at 50K/s for 30 minutes using JLBH to test Chronicle-FIX. The goal is to compare a system which is doing redundant work (in this case logging each message using SLF4J), compared with not logging (Chronicle-FIX records every message internally using Chronicle Queue) and how this changes the choice of Garbage Collector
- hinkley 2mo agoIt's like caching, in kind but not in type. Once you add it, people will stop trying to be parsimonious with resources and just reach for the cache every time. They'll just lean into it. In a hot minute you will discover you can't turn it off because people have lost their brains and the data flow of the app is through the cache and not through the call tree. If you tune for allocation patterns that are in the code, then you are cementing those as continuing in perpetuity. Better to cut the fat first, so that you can tune for the necessary complexity instead of the accidental. That will be self-correcting because any new misuses will be taxed with higher performance regressions.
- senderista 2mo agoI worked on Java code at AWS for a few years and nobody tried to optimize allocations. Then I changed jobs and started working on a Java MPP database and my first code review was brutal. You were expected to avoid allocations as much as possible (mostly by using the SoA pattern everywhere). At that scale no GC could save you from excessive allocations.
- motoboi 2mo agoI believe the whole string vs stringbuffer that later was made redundant by compiler contributed to that vision. People started dismissing allocation discipline as a thing from the past because "that thing was solved a lot ago and the compiler now is smart enough". Well, for string, yes, but not for arbitrary objects.
- cogman10 2mo agoThe most surprising allocation pressure I constantly run into is primitive boxing. The JVM does heroics to try and avoid it as much as possible, but when you end up with some primitive boxing in a hotspot the amount of GC pressure that creates can be unreal.
- drob518 2mo agoYep, and sometimes just a small code change can flip on boxing.
- cogman10 2mo agoOne of my least proud (most proud?) hacks when working with very large data sets is something like this Map<Integer, Integer> intCache = new HashMap<>(); while (loading) { Integer feild1 = intCache.computeIfAbsent(getField1(), (i)->i); } This is a terrible thing that shouldn't be as useful as it is to us... but it is really useful. We have a bunch of objects that can optionally have Integer values (hence a null is valid) but those int values are frequently the same. This saves a bunch of memory and ultimately GC pressure as a result. Valhalla can't come soon enough for us.
- drob518 2mo agoI do something similar for Java Time local dates. Financial data in particular has lots of redundant date info and benefits from being memoized. Converting to epoch millis also works.
- actionfromafar 2mo agoIf I squint, is this a special kind of heap compression?
- cogman10 2mo agoNot really heap compression or special, it's just reusing a reference to an object already allocated on the heap. Right now, if I do this LocalDate a = LocalDate.of(2020, 1, 1); LocalDate b = LocalDate.of(2020, 1, 1); A and B reference 2 different object allocations on the heap even though they are the same date. a != b. In Java, that can be pretty expensive even for an object as light as a LocalDate. By running the cache and doing var cache = new HashMap<LocalDate, LocalDate>(); LocalDate a = cache.computeIfAbsent(LocalDate.of(2020, 1, 1), (i)->i); LocalDate b = cache.computeIfAbsent(LocalDate.of(2020, 1, 1), (i)->i); Now you have the situation where `a == b` and you immediately end up dropping the object allocation for b on the next GC. The technique works best when you have a lot of repeated objects which are immutable. It is also only really needed because Valhalla isn't here. Once "value types" become a thing, then the representation for `LocalDate` inside the JVM can become just the fields and not a reference. The JVM is also free to do the sort of de-duplication optimization all on it's own for larger objects.
- zaphirplane 2mo agoHow does SoA help ?
- einpoklum 2mo agoGarbage collector? To quote Bjarne Stroustrup: > I don't like garbage. I don't like littering. My ideal is to eliminate the need for a garbage collector by not producing any garbage. That is now possible.
- bigfishrunning 2mo ago[flagged]
- drob518 2mo agoI prefer to see C++ as a failed experiment that just keeps going and going rather than garbage. The software industry learned a lot from it, both good and bad. But yea, I haven’t programmed in it since the late 1990s.
- scj 2mo agoI'd phrase it differently, C++ was a set of power-to-performance trade-offs that were optimal in the 1990s. Time has moved on. More importantly, a typical 1990s C++ dev was likely someone who learned assembly, then C or C++. Meaning they already knew how to control hardware / memory allocation, and C++ was just a new set of abstraction tools. It was a step forward for them. To modern devs, C++ is a step backwards. And a tough one at that.
- einpoklum 2mo agoActually, C++ was rather poor in the 1990s if you ask me (albeit still very usable). Time has moved on - but so has the language. Its implementation tradeoffs were much better IMNSHO after 2011; but it wasn't there yet. And it still isn't! It has a lot of warts that have to stay for backwards compatibility (which is a design goal); and then, it has annoyances I can't believe are not yet addressed (like - where is my 'restrict' keyword, damn it?!) Anyway, your view of the 1990s devs is incorrect. Almost no programmers who took up C++ learned assembly first (and few ever learned assembly). I believe most of them learned Pascal, or C, or scripting languages like Perl or Tcl or Unix shell scripts. Some may have learned Lisp or some ML variant as their first language, or Fortran 90. They didn't take a step back, they switched to a different set of language design goals and tradeoffs, and found it, well, serviceable. It is indeed a bit peculiar that the language has had this much staying power. For C, it's much more understandable - because C is such a small and simple language (and one which, as you suggested, often feels like a bunch of syntactic sugar over PDP-7 assembly). But C++ is big, and has its baggage and warts and flaws. I think it's probably because it's been able to adapt and stretch just enough under the influence of trends in programming languages, for people not to ditch it for something new. Maybe Rust will change that; but - C++ might very well "eat its lunch".
- peterabbitcook 2mo agoReading this gives me considerable pause - I can’t think of many classes within the codebase I work on that don’t have @Slf4j at the top… Since there wasn’t a link to the source code in that post, can you help me understand this - for the SLF4J baseline is your logger impl a console appender, a file appender, or a network service like an OTel collector? Does any of that matter for GC context?
- layer8 2mo agoAll common logging backends create a LogEvent or similar object for each logging call, and logging calls also typically construct new strings, which usually means a new StringBuilder object, its internal array (multiple ones if it grows), the final array it is copied to, and the String object that wraps that array. These are typically short-lived objects and therefore cheap. Nevertheless, continually creating many such objects increases GC pressure, in particular if the logging happens in code that doesn't otherwise create many objects.
- well_ackshually 2mo ago> All common logging backends create a LogEvent or similar object for each logging call, and logging calls also typically construct new strings, which usually means a new StringBuilder object, its internal array (multiple ones if it grows), the final array it is copied to, and the String object that wraps that array. Which then gets discarded because that was a Log.verbose and your minimum log level in production is WARN. Which is why many libraries have moved towards making your log message returned by a lambda. One constant lambda allocation (so, not a lot, an invokedynamic is absolutely fuck all.) that allows you to straight up skip allocating a full string that most likely is interpolating things and attempting to reach for context present on other threads is strictly better in 99.9% of the cases. The GC pressure is kept minimal and most importantly, constant.
- layer8 2mo ago> Which then gets discarded because that was a Log.verbose and your minimum log level in production is WARN. This isn't true for the LogEvent or equivalent object, which only gets created after the log level is tested to be applicable by the logger implementation. For call-site object allocation, you can wrap the logging call into an if statement that checks for the corresponding log level. The lambda allocation isn't constant if it captures anything from the surrounding scope, which will generally be the case for logging calls. (Unless by "constant" you mean that it's a single allocation per execution.)
- syngrog66 2mo agoI stopped reading early when it became clear the author didn't understand the terms (or perhaps the entire language) they were using. Several indicia.