7 ms·
SQLite: Wal2 Mode Notes
- ikawe 4y agoSQLite has different “modes” to facilitate recovery, this new WAL2 mode addresses the problem in WAL (version 1) where the recovery log could potentially grow very large. It’s solving a real problem, but not considered stable for production yet.
- jessermeyer 4y agoCitation needed.
- zimpenfish 4y agoFor which claim? The recovery log growing large is explained in [1] - "it does mean that the wal file may grow indefinitely [...] There are also circumstances [...] causing the wal file to grow indefinitely in a busy system." Not ready for production is [2] - although that is from 2 years ago, the WAL2 branch still isn't merged to trunk [3] which you'd expect if it was ready to go, I think? [1] https://www.sqlite.org/cgi/src/doc/wal2/doc/wal2.md https://www.sqlite.org/cgi/src/doc/wal2/doc/wal2.md [2] https://sqlite.org/forum/forumpost/17249fb83a?t=c&unf https://sqlite.org/forum/forumpost/17249fb83a?t=c&unf [3] https://www.sqlite.org/cgi/src/timeline?r=wal2 https://www.sqlite.org/cgi/src/timeline?r=wal2
- bob1029 4y agoI am curious what the situation is where the recovery log grows to large size and what the actual consequence of this would be. We have been using SQLite in WAL mode for over half a decade and never witnessed this. Several of our databases can see concurrent access from hundreds of users with transactions in the 1-10 megabyte range, so I find it a bit odd this never came up.
- zimpenfish 4y ago> I am curious what the situation is where the recovery log grows to large size From the link, "if a writer writes to the database while a checkpoint is ongoing [...] it does mean that the wal file may grow indefinitely if the checkpointer never gets a chance to finish without a writer appending to the wal file. There are also circumstances in which long-running readers may prevent a checkpointer from checkpointing the entire wal file - also causing the wal file to grow indefinitely in a busy system."
- muttled 4y agoI ran into this when I was importing about 1TB into a DB while simultaneously reading from it and performing tasks. This was all in a dev environment, but it did come up.
- Seattle3503 4y ago> what the actual consequence of this would be. SQLite runs in a surprising number of places. In an embedded environment disk space may be limited.
- pkhuong 4y agoI've seen the WAL grow when there's a constant write load and a long background read transaction keeps an old version alive. If you also never reach a point where there are 0 connections to the DB, you can keep this large WAL file around (most of it is unused) for a long time. I've had to restart services when WAL files had grown to multiple GBs and wouldn't shrink.
- eis 4y agoI wonder why they limit it to 2 WAL files instead of creating a new WAL file once the latest one reaches a certain size and a checkpoint is done. Then delete the WAL files that have been successfully flushed to the main DB. That could minimize the spikyness of WAL file garbage collection under less than optimal circumstances.
- samatman 4y agoAs an informed guess: backward compatibility. This way an SQLite library which has never heard of a WAL2 file will just use WAL mode instead, instead of potentially being confused by the unheard-of existence of multiple .wal files.
- ok_dad 4y agoI don't think that would work; there is also a -wal2 file which might be the active WAL file at the time a DB was closed. If you took an older library and tried to open that same DB, you would lose the data in the -wal2 file since that older library would ignore it.
- samatman 4y agoYou won't have a wal2 unless one of these connections generated it. The question is, can sqlite+wal2 write to a database while sqlite+wal is writing to it, without terrible things happening. The answer kinda has to be 'yes' right? This is SQLite we're speculating about here.
- ikawe 4y agoI think the idea is you only need 2 to achieve this new property of “avoiding unbounded growth”. You need the “one that is being written to” and the “other one which can be checkpointed into the main Db”. Reading between the lines a bit, I suspect unbounded growth is still possible with WAL2 - If you have an infinitely long running read transaction, data that modifies the page that’s being read from can never be checkpointed into the main DB. So setting aside the case of readers that never progress, with WAL1, even if your read transactions were finishing and fast forwarding through commits, it was possible to be in situations where, for a very long time (even potentially forever), at least one reader was behind the most recent write transaction, which meant the wal could never be truncated. Now with this WAL2 mode I think it should be guaranteed that so long as your read transactions are eventually ending and progressing towards more recent commits, you’ll always eventually find a time to checkpoint and truncate your WAL2 files.