7 ms·
I start to see a lot of these re-writes that depend on tests to state that its working. But the things that make software like Postgres and SQLite reliable are
by gingersnap 2mo ago
I start to see a lot of these re-writes that depend on tests to state that its working. But the things that make software like Postgres and SQLite reliable are not mostly the test, but the real world production scars. That's where the reliability comes from, years and years of running in production.
- kstrauser 2mo agoIn a project like PostgreSQL, those scars are reflected in unit tests demonstrating that they’re fixed. It’d be hard to pass its test suite and not be as robust as the original.
- dwedge 2mo agoSure but these scars/tests are from the original implementation. Just because it doesn't have issues there doesn't mean it didn't bring its own set of issues
- kelnos 2mo agoPassing a regression test suite only proves that those particular regressions aren't present. It proves nothing about robustness beyond that.
- ShinTakuya 2mo agoThis is all well and good in theory, but the number of times I've seen tests that don't actually test what they say they're testing is hard to count. Yes even when you encourage the developers to ensure the test fails first and do TDD. Tests help you ship with confidence but there's usually at least a few that are just passing by pure luck. So no, I wouldn't judge a rewrite as being equal just because it passes the tests. That said, I don't think that means you shouldn't do it. You just have to be pragmatic about it.
- simiones 2mo ago> It’d be hard to pass its test suite and not be as robust as the original. This is not true, even in principle, even for Postgres itself. You'd be right to say that it'd be hard to pass the test suite and not be robust at all to some extent. But even in Postgres, I bet that you can quite easily introduce a change that will pass the whole test suite but reduce robustness compared to the latest release (for a somewhat silly example, add a call to `exit()` on a timer that's longer than the longest duration test in the suite - that will significantly reduce robustness while still passing the entire test suite).
- guenthert 2mo agoThey ought to, but are they? In https://wiki.postgresql.org/wiki/Developer_FAQ https://wiki.postgresql.org/wiki/Developer_FAQ I don't see a requirement to provide a regression test for a bug fix.
- joshka 2mo agoIt would be reasonably easy to audit and automate this...
- jeltz 2mo agoIt is not a requirement the PostgreSQL project wants to have. It would be a heavy burden and mostly pointless.
- joshka 2mo agoIf you examine what you're saying here and slippery slope it a little, examining past effort for misses and correcting them at scale is not worth doing? There are many ways that the second part of your perspective are incorrect (burden / lack of point). I bet that a coding agent could in an automated fasion find at least reasonable additions that you'd say add value and reduce potential for error long term that you'd find valuable. They've probably already done so (I don't know postgres dev at all, so just supposition here. They will 100% do so in the future.
- guenthert 2mo agoI think it's rather unfortunate that writing regression tests is seen as a heavy burden. Compared to determining the cause of a bug and the right strategy to fixing it in a maintainable way, writing such a test should be simple, no? I'm not at ease regarding LLM generated code changes to a project with a (hopefully) long expected life time, but LLM generated regression tests should be less contentious. I wouldn't expect them to be maintained much; rather, if they don't perform as intended against a future build, just have them recreated.
- oblio 2mo agoEdsger W. Dijkstra: "Program testing can be used to show the presence of bugs, but never to show their absence!"
- tpetry 2mo agoYou immply that a testcase exists for every weird edge case. Especially filesystem and concurrency is things you can barely build test cases for. Even a 100% test coversge is far away from verifying all behaviour.
- kjs3 2mo agoI dunno...I can envision something vibecoded prioritizing passing test suites producing something that does that, but isn't even functional in real-world production. Sort of like in the pre-AI world, where someone claims 'standards compliance' by way of passing compliance test suites, but can't actually interoperate well with other implementations of the standard. YMMV.
- surajrmal 2mo agoUnit tests aren't useful for rewrites, only integration tests are. So there may be missing coverage. Also many things are simply difficult to test (eg performance under very specific conditions)
- jeltz 2mo agoPostgreSQL's tests are mostly integration tests.
- mrklol 2mo agoAnd also the amount of people running it in thousands of scenarios. Not sure if these areas can be even tested for, but I guess time will tell (can observe Bun if it breaks somewhere as that’s afaik the first big AI rewrite which got into prod for masses).
- joshka 2mo agoA lot of the signal (github, forums, mailing lists, discord, etc.) can be turned into signal. Right now it's easy enough to collect. In future it will be easy enough to cluster and generate preferences, experience, etc. Every bug report, code change as a result, PR / commit message, PR comment that steers preferences, etc. is solid signal to generate future tests.
- sshine 2mo ago> not mostly the test, but the real world production scars Most extensive test suites are exactly production scars: every time you have a bug or a regression, you write a test that confirms correct behaviour. SQLite is a good example to bring up because its extensive closed-source tests are what’s often cited as being what keeps people from forking it. (Turso did it, though, but it takes a company to deliver some guarantee of equivalent diligence.) And yes, years and years of running.
- hvb2 2mo agoThe maintainers that wrote those tests will have experience you won't get out of a rewrite. I think this is also where the real work is. A rewrite is one thing, that you can show off with a flashy blogpost. The maintenance, for years to come, won't be of that nature yet it still requires as much work.
- rustyhancock 2mo agoOne issue is those are the bugs you get when you write it in C++. They aren't the bugs you get when you write it in Rust. The kind of bugs you get are usually a function of the problem, language, implementation approach.
- consp 2mo agoSo you get other bugs when rewriting in another language without existing tests, got it. This is why I hate all the announcements of "it is rewritten in rust so it is obviously better than the original since it passes all the tests". Edit: and it's an LLM rewrite. Add that to the pile of over hyped messaging.
- baranul 2mo agoUnfortunately, too many people are getting captured by marketing and are divorcing themselves from reality. A rewrite can be an improvement, even if in the same or any other language. But, there are also levels, in terms of quality and human code review, when dealing with rewrites. New bugs can be introduced or there can be style issues, that can take time to fully reveal themselves, and particularly if the person or people involved are not familiar with the other language.
- rowanG077 2mo agoThat's precisely what a regression test suite is for. There is a bug, you fix the bug, you add a regression test. So if the test suite is well maintained these real world production scars are reflected in the tests.
- hk__2 2mo agoThe test suite is the result of these years of years of running in production. Every time you fix a bug, you add a non-regression test to ensure you don’t break it again.
- throwaway132448 2mo agoWait - does the AI rewrite the tests too? If so, lol.
- thunderbong 2mo agoI agree. I also agree with the sibling reply that - > every time you have a bug or a regression, you write a test that confirms correct behaviour. What I fail to see in these rewrites however is - what about new bugs introduced by virtue of this rewrite? I mean it'll have to go through its own challenges in real-world scenarios, right?
- deleted 2mo ago[deleted]
- zsoltkacsandi 2mo agoCompletely agree with this. The biggest lie of software engineering is that everything can be testable with tests. That a 100% test coverage is an indicator of quality software.
- Lomlioto 2mo agoI hope you are not true at all. Software like a Database should have an extensive test bench with concurrency tests, all corner cases etc. I'm not here running the new version on production to tell the maintainer/devs that my 'production unit tests failed'. What is this even for logic? I mean there is balance when i write tests for my production software, but my software is used by me. If i would have a library, i would test everything. And there was some blog post about another database system were they even virtualized the File access to test cases like when the disk controller stops working.
- xlii 2mo agoAs sibling mentioned - bugs and regressions are the thing that are (in a perfect world) usually covered. The problem however is non-covered success cases. A visualisation of the problem: let's say universe of interaction for DB consists of 10.000 SQL queries. Over 10 years various regressions were found and 2.000 SQL queries are guarded by tests. In reference implementation remaining 8.000 never surfaced over this time and it's unclear if they will work. And, thinking of how many various SQL queries PostgreSQL users around the world are using vs the test cases covered it's obvious that feature space isn't covered in 1% of the success ratio cases. Now the new, test-based implementation, has to prove it can handle remaining 99%.
- nextaccountic 2mo ago> I start to see a lot of these re-writes that depend on tests to state that its working. There's another way to validate the rewrite though. Just run both pgrust and postgres and compare the output. Know of an edge case? Run it too. Doesn't know? Use a fuzzer or some automated tool to find interesting inputs. Found an inconsistency? The input/output pair becomes a test case now Not sure if there's tooling for that though. If there is, just give it to Claude so they will incorporate it in their development loop
- booksock 2mo ago(I'm working with malisper on this) we built this too and are using it for a new version we're working on right now
- nextaccountic 2mo agoOh, that's cool!
- 27183 2mo ago> Not sure if there's tooling for that though. I can recommend proptest. What you're describing is a common pattern in property-based testing which basically boils down to "comparing against an oracle". In this case, postgres would be the oracle, pgrust is the system under test, and the idea is to generate strategies comprised of sequences of valid (and invalid) SQL statements and ensure the system under test behaves the same as the oracle in every case.
- mahogany 2mo ago> Just run both pgrust and postgres and compare the output. The space of inputs and outputs is infinite. You can't prove programs are the same by "just" testing a bunch inputs.
- nextaccountic 2mo agoIndeed. This approach is an improvement / augmentation over tests, not a panacea. Neither postgres nor pgrust have their behavior specified using formal methods. (pgrust could write some contracts using something like kani or creusot, but having upstream postgres also write contracts is a tougher sell). If they had, one could write a giant proof that said the two software essentially do the same thing (at least in a subset of environments and some simplifying assumptions)
- Suzuran 2mo agoThat's not relevant though. All concerns are secondary to security and Rust is the only language with security GUARANTEES. No other language is as secure. Therefore, even the worst Rust rewrite is automatically better than the best work in any other language, because it is the only one with guaranteed security. If a Rust rewrite of any of your software becomes available and you aren't installing it immediately and without reservation, then you are simply not giving security the priority it both demands and deserves, and that makes you disastrously insecure. This is a serious issue that should be given all priority. There is no room for debate. Your only policies should be security before all else and compliance with those policies must be absolute and without deviation, or all is lost.
- bbg2401 2mo ago> If a Rust rewrite of any of your software becomes available and you aren't installing it immediately and without reservation This is silly. Rust is awesome, and it's hard to argue against in many domains. However, software is more than the language it is written in or the runtime serving it. Is the Rust rewrite fully compatible? Is it supported by a strong community? Is it likely to continue to be supported? Is its release cadence sensible? Is its licence compatible with your intended usage? There are many questions needing to be answered before making rash decisions based purely on tech.
- Suzuran 2mo agoNone of those concerns approach the level of priority that must be assigned to security. On defense, your security must be perfect forever or you are absolutely defeated. None of us are on the red team. It's not a rash decision, it's the only decision that logic allows. If you are not secure, you are NOTHING.
- jeltz 2mo agoI strongly disagree. The easiest way to shut down your business is to insist on being 100% secure because the only way to have perfect security is to do nothing. Security is always about tradeoffs.
- dapperdrake 2mo ago"Everybody has a production system. The lucky ones also have a test system."
- alemanek 2mo agoNot weighing in on this specific rewrite but tests are how you specify that your software works correctly. If a behavior isn’t covered by an automated test in some form you can’t assert that any given change doesn’t break it. I think it is completely reasonable to use a preexisting unmodified test suite to state that something is working. The larger the project the more true this becomes. Real world production scars are documented and guarded against in the test suite otherwise those lessons get lost. Also SQLite is legendary for its massive test suite and extensive fuzzing. They have 590x the amount of test code and scripts than normal code. Source: https://sqlite.org/testing.html https://sqlite.org/testing.html
- LtWorf 2mo agoBut they aren't all open, so your llm rewrite/copyright eraser won't be taking advantage of them all.
- alemanek 2mo agoYeah I know. That is their moat. I was more responding to the assertion from OP that running in prod is what makes a product stable instead of tests. I was trying to call out that SQLite often credits their massive test suite for their stability but likely didn’t communicate that well.
- the__alchemist 2mo agoGreat point. Stated another way: "Mom, can I have battle-tested, reliable software" "We have battle-tested, reliable software at home" Battle-tested, reliable software at home: (Pic of green text from `cargo test`)
- edelbitter 2mo ago> That's where the reliability comes from So, we should make it easier to feed that reliability back upstream. Probably the most useful thing you can do with these LLM-transpilations for now: If the transpiled version passes all original tests, I can run my application test suite against it and use it to discover test coverage deficiencies in the original! If it crashes or otherwise observably misbehaves, I know the real project was missing regression tests for something. We could make upstream so much more resilient against accidentally breaking stuff in future updates, if only it becomes safe (offline + no side effects) and easy (if it crashes/locks, it is not from some memory safety bug from 25k transactions earlier) to run these transpiled projects as one row in our everyday integration matrix.