← All digests

DBMS Weekly — 2026-09-21 (week of Sep 21–27)

The week PostgreSQL 19 Beta 4 shipped, and the new features went through their first real stress test. Beta 4 was stamped on 21 September and announced on the 24th. It confirmed the reverts of SQL/PGQ, online checksums, FOR PORTION OF and MERGE/SPLIT PARTITIONS, and targets an RC in early October. The feature that survived, REPACK (CONCURRENTLY), took most of the testing. Thom Brown made it silently discard committed updates to a TOASTed column; Sami Imseih showed it could store NULL in a NOT NULL column (fixed and pushed on 25 Sep); Radim Marek found a hard catch-up ceiling at about 105M concurrent updates/deletes. The most important result came out of the same testing: Álvaro Herrera committed a fix, back-patched to 14, for initial logical-decoding snapshots that could set hint bits as if a committed transaction had aborted. Serializable isolation got the same treatment. Four separate threads showed SSI missing a conflict: TID range scans, TID scans, index-only scans racing heap_delete(), and the new ON CONFLICT DO SELECT. On the stable branches, 18.6 and 17.11 turned out to carry a standby-restart PANIC after visibility-map truncation. Outside the tree, PlanetScale wrote the clearest explanation yet of why vanilla Postgres lets you promote past unsynced logical slots, and Christophe Pettus took ClickHouse's physical-WAL replicator apart as a reimplementation of snapbuild.c.

PostgreSQL

  • What REPACK (CONCURRENTLY) costs while it runs — the benchmark the new feature needed: on a 90M-row, 27.8 GB table under 19beta4, REPACK matches VACUUM FULL on time (107 s) and writes less WAL than pg_repack (18.8 GB vs 33.5 GB). It also holds back the vacuum horizon cluster-wide through its slot, and it aborts once more than ~105M rows are updated or deleted during the run, because the combo-CID array keeps doubling. The author gives a rule of thumb, max (upd+del)/s ≈ 29,000 / REPACK hours, and reported the limit upstream (see mailing lists). (Radim Marek · boringsql.com)
  • The Shadow Knows — reads ClickHouse's physical-WAL→ClickHouse replicator as a from-scratch reimplementation of snapbuild.c/reorderbuffer.c. It still needs wal_level=logical and replica identity, it can silently diverge when an unchanged TOAST pointer is reused, and it ties your major-version upgrades to a vendor's audit of the WAL format. The proof offered is the 7 Sep varatt_external → varatt_external_oid rename: the physical format is not a contract. (Christophe Pettus · thebuild.com)
  • All your GUCs in a row: max_pred_locks_per_transaction and …per_page / per_relation — SSI internals with pgbench numbers. The default of 64 sizes an 8,704-target lock table for the whole cluster, and a separate 6,800-entry RWConflictPool has no GUC at all. On 18.6 at defaults, 75.6% of transactions failed with 40001; raising max_pred_locks_per_relation to 1000 cut that to 0.08% at about 4× the throughput. Raising the page threshold alone makes things worse. (Christophe Pettus · thebuild.com)
  • Our interleaved backfill wrote three times the WAL — backfilling 1M rows: range batches got 0% HOT and doubled the heap (269 → 548 MB). id % 10 = r interleaving at fillfactor 90 reached 96.1% HOT and kept the indexes flat. But with full_page_writes on and a checkpoint after each batch, interleaving dirtied every heap page every time: 3,022 MB of WAL vs 1,025 MB. Interleaving within an id range keeps HOT and caps WAL at 544 MB. Traced through heap_update/hio.c. (Mikhail Shytsko · seedfa.st)
  • Forty deletes vanished from pg_stat_statements — three reproduced blind spots that agent-generated SQL hits at once: the 5,000-entry cap evicts one-off statements; PG18 collapses two same-named tenant-schema tables into one queryid where PG17 kept them apart; and track=top hides everything inside PL/pgSQL and DO blocks. (Mikhail Shytsko · seedfa.st)
  • Blocking cutovers to save replication slots — why stock Postgres will promote a standby whose logical slots were never synced, which is a data-loss event for every CDC consumer. The fix is failover=true on the slot, sync_replication_slots=on and hot_standby_feedback=on, and the post explains why each defaults to off (they pin the vacuum horizon). (Simeon Griggs · planetscale.com) [vendor blog — substantive]
  • Anatomy of a (Postgres) search engine — the TIN implementers' own walk through inverted-index internals: term dictionary, postings compression, BM25 top-k block skipping, segment merges, tombstones. Then what Postgres specifically demands of one: a WAL-logged index relation, ctid mapping, VACUUM and visibility, and planner hookup through CustomScan. A general companion to last week's TIN post, not a product pitch. (Patrick Reynolds, Eric Ridge · planetscale.com) [vendor blog — substantive]
  • pgBackRest and PostgreSQL failover: why archive_mode matters — reproduces an incident where a standby was promoted with archive_mode=off. archive-mode-check=n does not bypass the check, and even a patched pgBackRest that forces the backup leaves the archive missing the timeline-switch segments and the .history file, so PITR across the promotion is impossible. Three-node cascading setup on PG18. (Stefan Fercot · pgstef.github.io)
  • Your Postgres database is slow, and it isn't Postgres — a field report: after a week of tuning shared_buffers and work_mem, the cause was that Azure disables host caching on managed disks of 4 TiB or more. The portal still lets you select "ReadOnly cache" on such a disk, and it has no effect. (Umair Shahid · stormatics.tech)
  • Ten years of Postgres logical replication — part one of three: a release-by-release table of what logical replication absorbed from Londiste, PgQ and pglogical (row/column filters in 15, partitions in 13, streaming and parallel apply in 14/16, origin=none in 16, conflict counters in 18, sequences in 19), then a hub-and-workers architecture rebuilt from core features only. (Dimitri Fontaine · tapoueh.org)
  • Being considerate of other people's time — a committer on what makes a patch cheap to review: finish the mechanical fixes before posting, label WIP with its known gaps, split big patches along the natural path (infrastructure → MVP with limits → relaxations) or by area. He also warns about vibe-coded patches whose authors can't explain them. It reads well next to this week's review-driven bug harvest. (Tomas Vondra · vondra.me) [by committer]
  • Breaking the Postgres superuser guardrails (part 2/6) — attacks on the hardening extensions that managed providers use to fake a restricted superuser. Techniques include CREATE FOREIGN DATA WRAPPER validator OIDs left unchecked during the superuser window, a $user search_path OID swap, and rewriting pg_proc.prosrc on a LANGUAGE internal function to bring back a blocked lo_export (→ RCE). The author reports 76 findings across providers. (Mehmet Ince · mehmetince.net) [unverified: provider-by-provider counts are the author's own]
  • Invoice numbers per tenant in PostgreSQL, without gaps — five approaches measured on 17.10 with 8 sessions, 1,000 tenants and 10% rollbacks. Only a per-tenant counter row updated with UPDATE … RETURNING gives zero gaps and zero duplicates. When the number is taken also matters: taking it last gives 746 invoices/s, taking it first gives 94/s, with the same guarantees. (Chris van Eijk · now-next.nl)
  • PostgreSQL 19 Beta 4 — (PGDG · postgresql.org)
    • reverted since Beta 3: SQL/PGQ property graphs, online enabling/disabling of data checksums, FOR PORTION OF temporal UPDATE/DELETE, ALTER TABLE MERGE/SPLIT PARTITIONS
    • fixes to REPACK, WAIT FOR, the RI fast-path FK check, autovacuum scoring and logical-replication conflict detection
    • RC planned for early October; GA "may also occur in October"
  • PgBouncer 1.26.0 — security release. (PgBouncer project · postgresql.org)
    • CVE-2026-19888: unauthenticated crash via a SCRAM client-final-message with no nonce (present since 1.11.0)
    • CVE-2026-6668: unauthenticated infinite loop from an integer overflow in packet-buffer growth
    • CVE-2026-6669: unbounded login work from a malicious server's SCRAM iteration count
    • also: tracks search_path by default, adds pool_idle_timeout

PostgreSQL mailing lists

  • [committers] Wait for transactions of an initial decoding snapshot to commit — the most important fix of the week, and it applies to every supported branch. SnapBuildInitialSnapshot() turned the builder's committed-xid list into an MVCC snapshot that could capture a transaction after its commit record but before its CLOG update, and then set hint bits as if it had aborted. That is real data corruption. The initial snapshot now waits for such transactions to leave the procarray. Found by stress-testing REPACK (CONCURRENTLY), but the bug is older than REPACK. (Álvaro Herrera; Antonin Houska, Rui Zhao; reported by Mihail Nikalayeu) [committed]
  • [hackers] REPACK (CONCURRENTLY) can silently lose updates when the TOAST table is rewritten — the decoding worker takes the TOAST relfilenode and drops its lock, but the backend doesn't lock the TOAST table until copy_table_data(). get_initial_snapshot() waits in between, so an ordinary open transaction is enough to open the window, and a VACUUM FULL of the TOAST table during it loses committed updates while verify_heapam() and bt_index_check() report nothing. Five patch versions in four days settled on Herrera's design of locking the TOAST relation up front in cluster_rel(). It is a PG19 open item. Its sibling bug, NULLs in NOT NULL columns added without a rewrite, was pushed on 25 Sep. The report that pg_dump can dump a table being repacked as empty was ruled expected behaviour, since REPACK (CONCURRENTLY) is not MVCC-safe until Antonin Houska's v20 series lands. (Thom Brown / shihao zhong / Álvaro Herrera · pgsql-hackers) [patch posted]
  • [hackers] REPACK (CONCURRENTLY) can't complete after ~105M concurrent updates/deletes — 75+ hostile runs found no showstopper, only a hard ceiling from combo CIDs during catch-up (about 106 MB per 2M replayed UPDATEs). The v20 fix replays changes with their original XIDs. For 19 it gets a doc note, and Antonin Houska argues the note should say that wraparound emergencies need failsafe VACUUM, not REPACK. (Radim Marek / Antonin Houska · pgsql-hackers) [patch posted]
  • [hackers] SSI can miss conflicts between index-only scans and heap writes — heap_delete()/heap_update() check for serializable conflicts before clearing the VM bit, on the assumption that the buffer lock keeps readers out. An index-only scan reads the old index entry while the VM bit is still set and never takes that lock, so two SERIALIZABLE transactions both commit. One of four SSI holes reported this week: TID range and TID scans allowing write skew (the fix takes a relation-level SIREAD lock, which reopens the false-positive trade-off cb5b28613d5 made in 2020), and ON CONFLICT DO SELECT returning a row covered by no SIREAD lock. (Rui Zhao; Andrey Borodin, Aleksander Alekseev, Zsolt Parragi, Dean Rasheed · pgsql-hackers) [patch posted]
  • [bugs] PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation — a regression in the current minor releases. A standby that has replayed a heap/VM truncation dies on restart with WAL contains references to invalid pages. 20 of 20 TAP runs fail on 18.6 and pass on 18.4 and 17.10; the original symptom was a three-node streaming cluster after a switchover. The follow-up says it doesn't need a switchover (with full_page_writes=off a plain restart hits it) and that a change already under review in Melanie Plageman's VM-clear thread fixes it. (Jacky Nguyen / Kirill Reshke · pgsql-bugs) [patch posted]
  • [committers] Avoid replay cleanup locks for freeze-only and VM-only records — since 17 combined pruning and freezing into one WAL record, redo has always asked for a cleanup lock, and 19 extended that to VM-only updates. Freezing moves no tuple storage, so standbys were waiting on buffer pins, and cancelling queries, for nothing. The fix is one line, back-patched to 17. (Melanie Plageman · pgsql-committers) [committed]
  • [bugs] BUG #19720: pg_trgm GiST index corruption from gtrgm_union() dropping SIGNKEY — when a union signature becomes all-true, the flag is set to ALLISTRUE without SIGNKEY. unionkey then treats it as an empty array, and a later insert writes a downlink carrying only the new value's bits, so searches can miss rows. Seen in production on 18.6; Kirill Reshke traced it to an oversight in 911e7020. (Nate Clark / Kirill Reshke · pgsql-bugs) [open]
  • [hackers] Up to 50x degradation in dblink performance when receiving notice traffic, 19 vs 18 — bisected to 112faf1378ee, which logs remote NOTICEs through ereport(). LOG passes the default log_min_error_statement, so every remote message repeated the full text of the running local query into the server log. With errhidestmt()/errhidecontext() the timings are back to v18. (Merlin Moncure / Tom Lane / vignesh C · pgsql-hackers) [patch posted]
  • [committers] Check EXECUTE privilege on functions invoked by the RI fast path — PG19's new FK-check fast path called the equality operator's function and any implicit cast without the EXECUTE checks the SPI path gets from ExecInitFunc(). Closed before GA, with tests showing the fast path and the partitioned SPI path now behave the same. (Amit Langote; reported by Nikolay Samokhvalov · pgsql-committers) [committed]
  • [hackers] Catversion bumps during beta — renaming pg_get_multixact_stats() before RC1 (it reports state, not stats) led Andrey Borodin to run an LLM review over every PG18→19 catalog addition for similar naming slips. Heikki Linnakangas then questioned the policy itself: avoiding late catversion bumps spares beta testers a pg_upgrade, and he calls that "a poor tradeoff". Also this week, master reverted the JSON_TABLE ON ERROR propagation over a pg_dump/pg_upgrade problem. (Heikki Linnakangas / Michael Paquier / Nathan Bossart · pgsql-hackers) [open]
  • [bugs] BUG #19722: window PARTITION BY numeric treats equal values with different scales as separate — the item is the reply: Tom Lane calls it "a near-duplicate of many recent reports" (five linked) and asks whether anyone has a real use case for testing scale() of a grouped numeric. In the same week, BUG #19712 arrived as an openly AI-written analysis of a multixact recovery deadlock already fixed in 15.19. Automated bug-finding now costs triage time as well as producing finds. (Tom Lane · pgsql-bugs) [open]
  • [hackers] Becoming a committer: form for tracking readiness — the committers are trying out a fillable form that summarises a candidate's most relevant work, for their own discussions of who to add. It is published so aspiring committers can fill it in themselves and plan against the long-standing criteria. (Noah Misch, on behalf of the committers · pgsql-hackers)

CommitFest (open: PG20-3, #62)

  • Balance (Sep 22 09:52 – Sep 27 18:37): ≥20 new · ≥5 closed (all committed) → net ≥ +15. (vs last week's ≥ +19: about −4. The per-CF activity log caps at 100 rows and its oldest row is 22 Sep 09:52, so all of Monday and Tuesday morning are clipped; both counts are lower bounds. One closure, "Reject WAIT FOR earlier…", is a second record for a patch already counted last week.) 8 entries moved to Ready for Committer. Queue at scan time: 90 needs review · 10 waiting on author · 19 ready for committer (119 active, up from 91).
  • New this week: Parameterized append subpaths (Alexander Pyhalov, Gleb Kashkin) — add_path() can discard a child's parameterized path that costs the same as an unparameterized path with pathkeys. With one child missing a parameterized path, no parameterized Append can be built, and the planner falls back to a hash join over foreign partitions. The choice then comes from a missing path, not a cost comparison: 1.7 s instead of ~2 ms in the repro. WIP proposal: keep all differently parameterized paths. Also new: Fix SSI conflicts between index-only scans and heap writes (#7344), Validate GIN posting lists before decoding them (#7352), Speed up lpad()/rpad(), BUG: pg_class.relchecks overflow (committed on 27 Sep as a 32,767 limit), Clear FatalError earlier during crash restart, pg_dump: fetch sequence data only for sequences being dumped.
  • Closed (all committed): REPACK (CONCURRENTLY) loses missing values of columns added without a rewrite (→ Herrera), Use bounded GIN pending-list cleanup in parallel autovacuum (→ Sawada), Assertion in pg_get_shmem_pagesize() in single-user mode (→ Linnakangas), DSA_ALLOC_NO_OOM vs dsm_create ERROR leaving a half-initialized pgstats hash entry (→ Paquier).

Community pulse

  • Is your Postgres migration safe or not safe? — a browser-only migration classifier from a former Cloudflare Postgres platform lead. The thread's main point: rule-based DDL checks can't be complete, because whether ALTER COLUMN TYPE is a no-op or a full rewrite depends on the current column type, not the statement. A commenter got ADD COLUMN status text NOT NULL with no default rated "safe". The most upvoted practical advice was about lock queuing, not DDL: an instant ADD COLUMN stuck behind a long read blocks everything behind it, and lock_timeout plus retry beat every tool. (Hacker News · 132 pts, 44 comments)
  • Postgres SELECT DISTINCT does not scale — DBOS's August post on DISTINCT walking every matching index entry, fixed with a recursive-CTE "loose scan", resurfaced. Commenters split between "DISTINCT is a code smell" and the real point: MySQL has loose index scan, PG18's skip scan doesn't cover this case, and the 2018 CommitFest attempt was abandoned after four years. The top structural reply: if the fix is a manual plan-forcing rewrite, it belongs in the planner. (Hacker News · 102 pts, 28 comments)
  • Postgres AT TIME ZONE 'UTC' does not do what you think it does — the second most upvoted reply corrected the article's opening premise: timestamptz stores no time zone either, since both types are the same 8 bytes. The useful model from the thread is that AT TIME ZONE converts in opposite directions depending on its input type, with the same syntax. Most agreed it's a footgun; a few said it's what you'd expect once the types are clear. (r/programming · 207 pts, 75 comments)
  • "Redis is a database?!" — the week's highest-comment database thread, a definitional argument started by an interview question. It sided against the poster: an in-memory cache is something you do with Redis, not what it is. The more interesting sub-thread tied the everyday meaning of "real database" to ACID, and asked where that leaves key-value stores that only promise eventual consistency. (r/Database · 283 pts, 310 comments)
  • Deleting 5 billion rows across 30 tables without drowning in bloat — PG17, a 150M-row delete from a table that is ⅓ heap, ⅓ TOAST, ⅓ index. Answers split between CTAS-and-swap in one transaction, batched deletes sized against the autovacuum scale factor and threshold, and pg_repack afterwards; one reply warned that partitioning is too late once you are at this point. (r/PostgreSQL · 21 pts, 28 comments)

Wider DBMS & distributed data

  • Postgres on NVMe: performance and the convergence of transactions and analytics — 8 identical m6id.4xlarge clusters on PG18.3 with pgbench at scale 33,000 (482 GiB, about 30× RAM). Local NVMe gave 9.2× the throughput of gp3 EBS (16,030 vs 1,734 TPS), 4.0 vs 36.9 ms per UPDATE, and VACUUM in 366 s vs 964 s. On EBS, 89% of wall time was spent off-CPU waiting on I/O, measured with Parca/eBPF and pg_stat_activity sampling. The second half is a CDC-to-ClickHouse pitch; the benchmark is the part worth reading. (Kaushik Iska · clickhouse.com) [vendor blog — substantive]
  • When to choose x86-64 vs aarch64 — Postgres-relevant ISA differences: a hyperthreaded x86 vCPU is half a core while a Graviton/Axion vCPU is a whole one, and AVX2/AVX-512 are 256/512 bits against Graviton's 128-bit vectors (which matters for pgvector distance kernels and bitmap intersection). It also covers the pg_trgm signed-vs-unsigned char corruption across architectures fixed in PG18, which is why indexes aren't byte-portable and an architecture switch means a logical-replication migration. (Ahmed Darwich · planetscale.com) [vendor blog — substantive]
  • What happens when the model eats the stack? — review of the Berkeley/Stanford data-agents paper: a plain coding agent on a frontier model beats hand-built data agents on accuracy and tokens (turns fall from 23.2 to 6.0), and over 60% of the remaining failures are semantic (schema, join keys, business definitions). The paper's claim is that persistent semantic context is the research problem that will last. (Murat Demirbas)
  • The Humanoid Lesson, or, will AI leave anything for database researchers to do? — a skeptical counter-review of the same paper: "persistent semantic context" is tribal knowledge, a decades-old problem, so keep the docs in git next to the code. His rule for predicting where an LLM struggles is to ask where a person would. (A. Jesse Jiryu Davis · emptysqua.re)

Commercial engines (SQL Server, Oracle, MySQL, …)

  • io_uring in Oracle Database: a hybrid storage I/O architecture at production scale — Oracle's own account of replacing libaio: per-process rings and shared buffer registration (now in mainline Linux), with libaio kept as a fallback. TPC-C throughput is unchanged with server CPU down 1.2 points, TPC-H CPU per query is down 8.5% (geomean), and the write path alone uses 29% less CPU per write with 34% more throughput. Synchronous reads gain nothing, and CPU rises for very large I/Os, so the conclusion is pread/pwrite for sync I/O and io_uring only for async batches. That maps directly onto Postgres's AIO design choices. (Chowdhury, Shah, Susairaj, Agrawal) [paper] (submitted 19 Sep, announced 22 Sep)
  • TideSQL 5 and MariaDB: the tides are moving fast — an LSM storage engine (TidesDB 10) for MariaDB, now with MVCC, secondary indexes, FKs, XA and Galera. The post reads the vendor's own sysbench critically: the "60× InnoDB" headline ran InnoDB with a 128 MB buffer pool on a 3 GB dataset, and with 4 GB caches InnoDB wins point selects by about 2.9× while TideSQL stays ahead on write-heavy runs. It also flags that optimistic-conflict retries surface as MariaDB error 1180, not the 1213 applications retry on. (Frédéric Descamps · mariadb.org)

Research & cutting edge

  • Batched feedback and the random-access wall in search-based graph construction — splits ANN graph build time into distance-evaluation count and cost per evaluation. A batch builder whose graph changes in B synchronous blocks does 0.88× Vamana's distance work but takes 1.55× the wall-clock time, and every search-based builder runs 20–27× below a dense kernel at d ≈ 100. The gap grows as d^0.4, and batching the beam search afterwards recovers only 2–3%. It explains why vector index builds are memory-bound. (Chávez) [paper]
  • Same pattern, different answer: GQL and SQL/PGQ path-pattern divergence — an executable reference semantics for the path-pattern core shared by GQL and SQL/PGQ, run as a 17-construct suite against 6 releases of 5 engines (Kuzu, DuckPGQ, Neo4j, Memgraph, Apache AGE). 9 of 17 constructs get more than one answer across engines, and 15 of the 26 disagreements are silent: no error, just a different result. Useful background for when SQL/PGQ comes back to Postgres. (Mandarapu, Kunkunuru) [paper] (submitted 19 Sep, announced 22 Sep)
  • Provenance of HAVING queries in semirings with monus — provenance for aggregate-condition queries in any commutative semiring with monus, with no extra operators, matching the provenance of the aggregation-free self-join rewrite of HAVING COUNT(*). Implemented in ProvSQL, a PostgreSQL extension. The abstract gives no numbers. (Sen, Karmakar, Maniu, Saadeh, Senellart) [paper]
  • Attacking Diophantus: special cases of bag containment — proves bag-semantics containment of conjunctive queries decidable when the contained query is "join-uniform", with the containing query arbitrary. That covers both previously known decidable cases, via a Diophantine inequality system over a canonical model built from all unifications of the contained query. The theory behind which bag-semantics rewrites are safe. (Konstantinidis, Li, Mogavero) [paper]
  • Pierce: GPU ray tracing for spatial joins over complex 3D data — recasts 3D polyhedral spatial joins as ray tracing on RT cores: rays cast along one mesh's edges against a hierarchy built over the other, with ray–node tests for pruning and ray–triangle tests for refinement. Reports "more than two orders of magnitude" over the state of the art on digital-pathology data. (Hackl, Zacharatou · SIGSPATIAL 2026) [paper]
  • QUIVER: dense vector search inside a SPARQL engine — native vector search in QLever. Parsing JSON vectors once at load time instead of per query gives median speedups of up to 41.9× (BSBM) and 20× (DBpedia); an ANN index exposed as a virtual SERVICE raises that to 355× and 97.8×. Cross-modal vector joins that time out everywhere else finish in seconds. The general lesson is where to parse and where to index a vector column. (Kantz, Schreck, Silvello) [paper]
  • Exploiting residual reachability for cross-model migration of graph-based ANN indexes — when the embedding model changes, build the new graph from the old one (shallow expansion with sign-code-screened second-hop candidates, or a hop-bounded beam search) instead of from scratch. Up to 17.43× faster than the fastest degree-matched rebuild across eight migrations, "while keeping competitive recalls". (Gu, Zhong, Jin, Cheng, Ni, Li, Song, Shen) [paper] (submitted 19 Sep, announced 22 Sep)
  • KathDB-FAO: synthesized query plans in a multimodal DBMS — compiles a natural-language query into atomic actions with input/output contracts, groups them, and synthesizes code per group at execution time instead of calling an LLM per tuple. 58.8% lower execution cost than the next best system on SemBench at equal or better quality. (Xiao, Brown, Borycki, Balazinska) [paper]

International (non-English sources)

  • Why an ultra-fast COMMIT is dangerous, and how to check that WAL really reaches disk — single-client pgbench shows COMMIT at 0.223 ms with a real fsync against 0.060 ms when the sync is skipped (~1,370 vs ~1,920 TPS per client). A known ~135 µs flush on an Intel P3700 shows 60 µs can't include one. Four checks follow: strace for missing fdatasync, iostat f/s stuck at 0, a FUSE layer that logs syncs, and a sysrq b crash test, which on a failing system loses committed rows and gives invalid magic number 0000 on the replica. The system tested is never named. (slonik_pg · Postgres Professional on Habr) [ru] (orig: Почему сверхбыстрый COMMIT опасен для СУБД и как проверить надёжность сохранения данных)
  • 1C optimizer's notes, part 20: how much Huge Pages actually buy Postgres — PG17 on a Proxmox VM with 8 vCPU, 100 GB RAM and 40 GB shared_buffers, a 180M-row table prewarmed, pgbench with 8 clients, perf stat system-wide. Index lookups went 54,263 → 58,564 TPS (+7.93%) and page faults over 25 s fell from 1,397,551 to 1,044; a join/aggregate query gained 4.57%. 1 GB pages were slightly worse than 2 MB (57.6k vs 58.5k TPS). (gallam · SOFTPOINT on Habr) [ru] (orig: Записки оптимизатора 1С (ч.20). На сколько реально настройки Huge Pages для Postgres могут ускорить запросы 1С)
  • I explained an ora2pg bug wrong twice — a migration war story. ora2pg drops foreign keys when migrating MySQL/MSSQL schemas because a nested Perl exists() in create_unique_keys() autovivifies a fake partitions_list entry for every table with a PK/UNIQUE key. create_foreign_keys() then skips the FK whenever pg_version <= 12, and an unset PG_VERSION defaults to 11. Setting PG_VERSION ≥ 13 fixes it; AUTO_INCREMENT start values are only read from a live database. (Lunch418 · Habr) [ru] (orig: Я дважды неправильно объяснил один баг в ora2pg)

Upcoming events

  • PGDay Israel 2026 — Tel Aviv, 25 October — free one-day community event; the program is published. Internals picks:
    • REPACK In Core (Robert Treat) — arrives the week after REPACK (CONCURRENTLY)'s first stress test turned up lost updates, a catch-up ceiling and an MVCC-safety gap; worth hearing what the design promises against what the list found.
    • My experience with AI for Postgres core development (Andrey Borodin) — from someone who filed LLM-found SSI and catalog-naming issues this very week.
    • PostgreSQL I/O: from synchronous to modern asynchronous I/O (Lior Friedler) — read alongside Oracle's io_uring paper above: which I/O paths gain from async and which don't.

New sources added this week

  • pgstef.github.io — pgBackRest failure modes reproduced end to end on real multi-node setups. (Stefan Fercot)
  • now-next.nl/en/insights — a measured multi-tenant Postgres series (indexes, gapless numbering, RLS, partitioning), each post benchmarked on PG17. (Chris van Eijk)
  • emptysqua.re — sharp DB-research paper reviews and TLA+ writing; a natural companion to Murat Demirbas. (A. Jesse Jiryu Davis)
  • dbos.dev/blog — Postgres postmortems from running a durable-execution engine on it (DISTINCT scaling, MVCC delete locality); product mix, judge per post. (Peter Kraft, Qian Li)
  • SOFTPOINT on Habr — «Записки оптимизатора 1С» — 20+ parts of measured 1C-on-PostgreSQL work with perf counters and stated methodology.

58 items · yield — mailing lists: 926 messages in window (790 hackers / 116 bugs / 0 performance / 20 general, plus 133 pgsql-committers) → 22 shortlisted → 12 published · blogs: 46 posts in window → 39 shortlisted → 20 published · community: ~215 threads viewed → 10 shortlisted → 5 published · research: 32 cs.DB preprints announced in window (20 submitted in window) → 13 shortlisted → 9 published · international: ru 27→7→3, zh 1→1→0, ja 5→2→0, fr 1→1→0.