7 ms·
I disagree. I work in a large environment, and rsyslog is where 90% of data goes to first. It can keep up with millions of messages per second, route them to hi
by skullone 25d ago
I disagree. I work in a large environment, and rsyslog is where 90% of data goes to first. It can keep up with millions of messages per second, route them to higher order services for indexing (bigquery, splunk, elastic etc etc). Has rules engines, encryption, supports multiple protocols and obviously has TLS too. You can surely augment with otel and such where you can, but syslog is uhhhh, deployed in so many places that it would make an average app developer's head spin when all they're used to is application logging in a controlled structured place in their silo.
- shmoe 25d agoEven SC4S, splunk's docker appliance for turnkey syslog uses rsyslogd. Edit: being pedantic -- it's syslog-ng actually.
- skullone 25d agoAnd the number of k8s envs that log stdout through them into.... more rsyslog, it's truly everywhere. Plus all the sidecar containers deployed that shuffle app logs, lots of syslog there, its so lightweight and simple and reliable. I watch all the gyrations people go through to achieve the same result, and it's always changing, hurts my brain thinking how much time they waste
- otterley 25d agoOn the other hand, traditional syslog is UDP based, so as soon as the receiver experiences CPU or I/O starvation and its receive buffer overflows, it will begin dropping messages. That's not great for observability, and may well be impermissible at many sites that need end-to-end log integrity (e.g. audit logs).
- lanstin 25d agocheaply dropping log msgs you cannot handle is absolutely essential for an observability system - otherwise excess load can take down the logging infra which can (if msgs aren't dropped) take down the prod network/app trying to send reliable log msgs. Audit logs are a distinct feature.
- otterley 25d agoThat's fine if you plan for it and can clearly delineate which logs can be dropped and which can't. The challenge is that most applications log to a single stream that consists of both high and low-priority logs muxed together and is sent to a single destination, making it impossible to distinguish the two, and there's only one receive buffer. Excess load can't take down a well-engineered log collection infrastructure. There can be overload, but the backpressure should propagate downstream and senders and intermediaries should buffer locally if needed. Once the collectors are able to catch up again, the spooled messages will be dispatched, and the backlog should recover. A well-engineered logging system for sites that care about integrity and durability should look a lot like a distributed message queue. > Audit logs are a distinct feature. In my experience, this is not always as distinct as one might hope. On multiple occasions in my career, a customer demanded we perform research using our logs to answer, and the information they sought were not in the class of logs that were considered "audit logs" in advance. Everyone chooses differently what qualifies as "audit logs"; it doesn't have an objective definition.
- skullone 25d agoIf you are doing massive log streams with mixed priority/durability, use named queues if you have different priority and retry and durability needs. You keep bringing up edge cases, but they've already been considered and solutions already engineered and available in every major syslog implementation.
- otterley 25d agoI didn't say they those choices aren't viable; I'm strictly talking about architectural decisions. You can solve the problem with different solutions, be they rsyslog or otherwise. That said, I probably wouldn't go with rsyslog as my default choice anymore since the world is moving on to OpenTelemetry.
- delamon 25d ago> There can be overload, but the backpressure should propagate downstream and senders and intermediaries should buffer locally if needed. You're only delaying the inevitable. Even with local buffering, you can arrive at a point where you can buffer no more and have either to choke the production workload or start dropping messages.
- skullone 25d agoYou can use TCP and other mechanisms for guaranteed delivery. Like, almost all LIDR logging across the planet uses TCP syslog, which is durable and attestable to in courts.
- otterley 25d ago> TCP syslog...which is durable and attestable to in courts. TCP alone won't get you there. It's certainly not durable in and of itself. All TCP can do is ensure that streamed data is received in the correct order, and confirm that a segment's data was successfully placed into the right buffer on the receiver side. You also need immutable storage, stronger integrity checks than what TCP itself provides, and many other requirements I haven't researched in a while.
- skullone 25d agoAs someone whose built megabyte to petabyte scale system of records, "yah but" to nits is annoying. You've solved it all yourself in your thought exercise though. Saying "the world has moved on", eh, not in any sense of the legal world, no. There's entire ecosystems around syslog alone to satisfy everything, otel is a baby fart feature and reliability and attestable wise
- deleted 25d ago[deleted]
- otterley 25d agoThere’s no need to be rude and dismissive. Knock it off. You’re coming across as unnecessarily provocative and defensive. If you truly believe OTel is a “baby fart,” I’d recommend you go to SRECon, KubeCon, and other similar conferences and make your opinion loudly known. Nobody talks about syslog there, and these are big and well-respected companies represented there. Syslog is mature and battle tested for what it is: schlepping unstructured plaintext logs from one place to another. It’s not really purpose-built to meet the needs of a full-fledged observability solution, though, of which logs are but a component.
- xorcist 25d agoThirty years ago and more that might have been a valid objection, but at that time the alternatives weren't great either. No one has suggested running syslog over unreliable transport after that. In fact, the queue management and at least the possibility of some rudimentary end-to-end cryptographic integrity checks are some of the stronger points of rsyslog. Splunk Cloud and Elastic, as far as I know, lacks the latter completely which rules them out as a single log sink for environments with that type of requirements.
- otterley 25d agoWhich major Linux distro ships rsyslog with TCP as the default remote protocol and durable local-buffer configuration out of the box for remote delivery? Genuinely curious. A modicum of research reveals that even the rsyslog documentation starts out with UDP for remote delivery: https://docs.rsyslog.com/doc/getting_started/beginner_tutorials/06-remote-server.html#configure-the-server-receiver https://docs.rsyslog.com/doc/getting_started/beginner_tutori...
- skullone 25d ago"Beginner tutorial", I'm sorry for not taking your point seriously, but I can't take it seriously. TCP for syslog (and RELP) have been around a long time (late 90s for syslog over TCP, 2006 for RELP). rsyslog and syslog-ng support it all, and operators have had choices given the import of their log data and what they can tolerate.
- otterley 25d ago> "Beginner tutorial", I'm sorry for not taking your point seriously, but I can't take it seriously. Well, maybe go observe how a broad array of sites implement it in practice, then you might take it more seriously. Maybe you don't implement it that way, but a lot of people will just follow the tutorials or shortcut their way to something that works (but is brittle). At any rate, I was responding directly to the claim that "No one has suggested running syslog over unreliable transport" which is obviously untrue.
- firesteelrain 25d agoSyslog supports TCP and UDP? We run dual Syslog servers and support both protocols
- Bender 25d agotraditional syslog is UDP based This was a solved problem a long time ago in rsyslog. One can define a local spool and enable TCP (and optionally encryption) to multiple syslog servers. If something interrupts the flow the syslog messages will queue locally and then de-spool when communications are restored.
- otterley 25d agoYes, we’ve discussed that in multiple threads. But it’s not the default and it’s unclear how often this is used in the wild. https://news.ycombinator.com/item?id=49426950 https://news.ycombinator.com/item?id=49426950
- firesteelrain 25d agoIn our case, other than network switches and routers, Syslog is the log of record and Splunk runs on all the VMs. Splunk isn’t forwarding syslog itself