| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
PDU (PostgreSQL Data Unloader) is a professional-grade, open-source disaster recovery and data extraction tool for PostgreSQL databases (versions 14-18). It reads PostgreSQL data files directly without requiring a running database instance—making it the ultimate solution for disaster recovery, forensic analysis, and emergency data extraction.
PostgreSQL DBAs face critical scenarios like these:
▸ Database completely corrupted - won't start, can't recover ▸ Accidental DELETE/UPDATE - critical data gone in seconds ▸ Data files deleted - filesystem corruption or operator error
Traditional tools like pg_filedump, pg_dirtyread, and pg_waldump each solve one problem with different learning curves. PDU unifies all these scenarios into a single, intuitive tool.
💬 In Oracle ecosystems, professionals use ODU/DUL for direct data mining. PostgreSQL deserved the same level of capability—that's why PDU exists.
|
Two files (pdu + pdu.ini), zero dependencies. One tool handles corruption, deletion, and updates. |
100% read-only operations. Your original data files remain untouched. Recover data faster than traditional PITR methods. |
Powerful direct data file reading capabilities. Supports PostgreSQL 14-18. |
PDU handles three critical disaster scenarios with two recovery methods:
| Scenario | What Happened | PDU Solution | Guide |
|---|---|---|---|
| ▸ Database Corruption | PostgreSQL won't start due to catalog/file corruption | Extract data directly from damaged files | Read → |
| ▸ Accidental DELETE/UPDATE | Wrong DELETE/UPDATE executed in production | Replay WAL archives to recover original data | Read → |
| ▸ Data File Loss | Specific data files deleted or corrupted | Rebuild data from remaining files | Read → |
1. Direct Data File Extraction
2. WAL Archive Scanning
# Set PostgreSQL version (14-18)
sed -i 's/#define PG_VERSION_NUM [0-9]\+/#define PG_VERSION_NUM <Version>/g' basic.h
# Compile
make
# Result: ./pdu executablePGDATA=/var/lib/postgresql/15/main
ARCHIVE_DEST=/var/lib/postgresql/wal_archive./pduPDU> b; -- Bootstrap metadata from PGDATA
PDU> \l; -- List databases
PDU> use production_db; -- Switch database
PDU> \dn; -- List schemas
PDU> set public; -- Switch schema
PDU> \dt; -- List tables
PDU> \d+ customers; -- Describe table structure
PDU> unload tab customers; -- Export one table to CSV📖 Full Guide: pduzc.com/quickstart
Problem: PostgreSQL won't start due to corrupted system catalog.
Solution:
PDU> b; -- Initialize metadata
PDU> use mydb;
PDU> set public;
PDU> unload tab customers; -- Export one table
PDU> unload sch public; -- Export all tables in schema
PDU> unload ddl; -- Generate DDL for current schema
PDU> unload copy; -- Generate COPY script for exported CSVsResult: CSV files exported to current directory, even if PostgreSQL won't start.
Problem: A critical DELETE statement was executed, and you need the deleted rows back.
Solution:
PDU> use production;
PDU> set public;
PDU> param restype delete; -- Scan DELETE records
PDU> param resmode tx; -- Restore by transaction ID
PDU> scan orders; -- Scan WAL for deleted rows
PDU> restore del <TxID>; -- Restore one transaction from scan resultsResult: Deleted rows recovered from WAL archives and exported to CSV.
For time-range recovery, use param resmode time;, set param starttime <YYYY-MM-DD HH:MM:SS>; and param endtime <YYYY-MM-DD HH:MM:SS>;, then run scan orders; and restore del all;.
Problem: An UPDATE statement modified hundreds of rows with wrong values.
Solution:
PDU> use production;
PDU> set public;
PDU> param restype update; -- Scan UPDATE records
PDU> param resmode tx; -- Restore by transaction ID
PDU> scan users;
PDU> restore upd <TxID>; -- Restore pre-UPDATE valuesResult: Original values before UPDATE statement recovered from WAL.
📖 More Scenarios: pduzc.com/docs/instant-recovery
| Command | Description |
|---|---|
| b; | Bootstrap metadata from PGDATA |
| use <database>; | Switch to specified database |
| set <schema>; | Switch to specified schema |
| \l; | List all databases |
| \dn; | List schemas in current database |
| \dt; | List tables in current schema |
| \d+ <table>; | Describe table structure (columns, types) |
| \d <table>; | Show table column types |
| u tab <table>; / unload tab <table>; | Export one table to CSV |
| u sch <schema>; / unload sch <schema>; | Export all tables in a schema |
| u ddl; / unload ddl; | Generate DDL for the current schema |
| u copy; / unload copy; | Generate COPY statements for exported CSVs |
| scan <table>; | Scan WAL archives for DELETE/UPDATE records of one table |
| scan manual; | Initialize metadata from files under the manual directory |
| scan drop; | Scan WAL for dropped/truncated table metadata |
| meta tab <table>; | Write one table's metadata to tab.config |
| meta sch <schema>; | Write all table metadata for a schema to tab.config |
| restore del <TxID>; | Restore deleted rows for one scanned transaction |
| restore upd <TxID>; | Restore pre-UPDATE rows for one scanned transaction |
| restore del all; / restore upd all; | Restore scanned results in time-range mode |
| add <filenode> <table> <columns>; | Manually add table metadata; put data files in restore/datafile |
| restore db <db> <path>; | Initialize a custom database directory (Pro/Enterprise) |
| dropscan idx; / ds idx; | Build dropped-page index (Pro/Enterprise) |
| dropscan; / ds; | Scan disk using tables in tab.config (Pro/Enterprise) |
| dropscan iso; / ds iso; | Recover from an image file using tables in tab.config (Pro/Enterprise) |
| dropscan repair; / ds repair; | Retry failed TOAST table scans (Pro/Enterprise) |
| dropscan clean; / ds clean; | Clean restore/dropscan directories (Pro/Enterprise) |
| dropscan copy; / ds copy; | Generate COPY commands for DropScan output (Pro/Enterprise) |
| p startwal <WAL>; / param startwal <WAL>; | Set WAL scan start file |
| p endwal <WAL>; / param endwal <WAL>; | Set WAL scan end file |
| p starttime <YYYY-MM-DD HH:MM:SS>; / param starttime <YYYY-MM-DD HH:MM:SS>; | Set time-range recovery start time |
| p endtime <YYYY-MM-DD HH:MM:SS>; / param endtime <YYYY-MM-DD HH:MM:SS>; | Set time-range recovery end time |
| p resmode tx|time; / param resmode tx|time; | Set recovery mode |
| p restype delete|update; / param restype delete|update; | Set recovery type |
| p exmode csv|sql; / param exmode csv|sql; | Set export format |
| p encoding utf8|gbk; / param encoding utf8|gbk; | Set output encoding |
| p isomode on|off; / param isomode on|off; | Set image-save mode |
| reset <parameter>; / reset all; | Reset one parameter or all parameters |
| show; | Display current parameters |
| t; | Display supported data types |
| exit; / \q; | Exit PDU |
All interactive commands must end with ;.
📖 Full Reference: pduzc.com/docs
| Parameter | Description | Example |
|---|---|---|
| PGDATA | PostgreSQL data directory path | /var/lib/postgresql/15/main |
| ARCHIVE_DEST | WAL archive directory (for DELETE/UPDATE recovery) | /var/lib/postgresql/wal_archive |
📚 Comprehensive guides available at pduzc.com/docs/instant-recovery
| Scenario | Description | Official Guide |
|---|---|---|
| Corrupted Instance | PostgreSQL won't start—extract data straight from disk | Open Guide → |
| Corrupted Datafile | Specific datafile damaged—salvage reachable data | Open Guide → |
| Corrupted Database | Entire database broken—reconstruct catalog and data | Open Guide → |
| Deleted Records | Rows accidentally deleted—replay WAL to recover | Open Guide → |
| Updated Records | Wrong UPDATE executed—rollback with WAL analysis | Open Guide → |
| No Backup: DELETE / DROP TABLE / DROP DATABASE | Compare PITR, pg_dirtyread, pg_filedump, PDU, and professional recovery | Decision Guide → |
Q: make fails with "lz4.h: No such file or directory"
A: Install LZ4 development library:
# Ubuntu/Debian
sudo apt-get install liblz4-dev zlib1g-dev
# RHEL/CentOS
sudo yum install lz4-devel zlib-develQ: How do I change the PostgreSQL version target?
A: Edit basic.h and modify PG_VERSION_NUM:
sed -i 's/#define PG_VERSION_NUM [0-9]\+/#define PG_VERSION_NUM 16/g' basic.h
make clean && makeQ: PDU shows "PGDATA directory not found"
A: Verify the PGDATA path in pdu.ini is correct and accessible:
ls -la /var/lib/postgresql/15/mainQ: WAL scanning (scan <table>) returns no results
A: Ensure:
Q: Exported CSV files contain garbled characters
A: Check database encoding:
PDU> \l; -- Look at the encoding columnThe CSV will use the database's encoding. Convert if needed:
iconv -f UTF8 -t GBK input.csv > output.csvQ: Unloading large tables takes too long
A: PDU performs full table scans. For better performance:
sudo apt-get update
sudo apt-get install build-essential liblz4-dev zlib1g-devsudo yum install gcc lz4-devel zlib-develWe welcome contributions from the community! Here's how you can help:
Found a bug or have a feature request? Open an issue
Help make PDU accessible to more communities by translating the documentation.
This project is licensed under the Apache License, Version 2.0. See LICENSE for the full license text.
Historical Note: Earlier drafts used the Business Source License 1.1. Starting from this release, PDU is 100% open-source under Apache-2.0 for maximum community compatibility.
This project includes code derived from:
See NOTICE for complete attribution.
ZhangChen (Jay) Copyright © 2024-2025
If PDU helped you recover critical data or saved your day, please consider:
⭐ Starring this repository on GitHub 📢 Sharing PDU with your team 🐛 Reporting issues to help improve the tool 💬 Joining discussions and sharing your use cases
Visit pduzc.com for comprehensive guides, or open an issue on GitHub.
Made with ❤️ for the PostgreSQL Community
PDU (PostgreSQL Data Unloader) 是一款专业级、开源的 PostgreSQL 数据库灾难恢复和数据提取工具,支持 PostgreSQL 14-18 版本。它可以直接读取 PostgreSQL 数据文件,无需运行中的数据库实例——是灾难恢复、取证分析和紧急数据提取的终极解决方案。
PostgreSQL DBA 经常面临这些关键场景:
▸ 数据库完全损坏 - 无法启动,无法恢复 ▸ 误执行 DELETE/UPDATE - 关键数据瞬间消失 ▸ 数据文件被删除 - 文件系统损坏或操作失误
传统工具如 pg_filedump、pg_dirtyread 和 pg_waldump 各自解决一个问题,且有不同的学习曲线。PDU 将所有这些场景统一到一个直观的工具中。
💬 在 Oracle 生态系统中,专业人士使用 ODU/DUL 进行直接数据挖掘。PostgreSQL 也应该拥有同等级别的能力——这就是 PDU 存在的原因。
|
两个文件(pdu + pdu.ini),零依赖。一个工具处理损坏、删除和更新的恢复。 |
100% 只读操作。您的原始数据文件保持不变。比传统 PITR 方法更快地恢复数据。 |
强大的直接数据文件读取能力。支持 PostgreSQL 14-18。 |
PDU 使用两种恢复方法处理三种关键灾难场景:
| 场景 | 发生了什么 | PDU 解决方案 | 指南 |
|---|---|---|---|
| ▸ 数据库损坏 | PostgreSQL 因目录/文件损坏无法启动 | 直接从损坏文件中提取数据 | 查看 → |
| ▸ 误执行 DELETE/UPDATE | 在生产环境执行了错误的 DELETE/UPDATE | 重放 WAL 归档以恢复原始数据 | 查看 → |
| ▸ 数据文件丢失 | 特定数据文件被删除或损坏 | 从剩余文件重建数据 | 查看 → |
1. 直接数据文件提取
2. WAL 归档扫描
# 设置 PostgreSQL 版本(14-18)
sed -i 's/#define PG_VERSION_NUM [0-9]\+/#define PG_VERSION_NUM 大版本号/g' basic.h
# 编译
make
# 结果:./pdu 可执行文件PGDATA=/var/lib/postgresql/15/main
ARCHIVE_DEST=/var/lib/postgresql/wal_archive./pduPDU> b; -- 从 PGDATA 引导元数据
PDU> \l; -- 列出数据库
PDU> use production_db; -- 切换数据库
PDU> \dn; -- 列出模式
PDU> set public; -- 切换模式
PDU> \dt; -- 列出表
PDU> \d+ customers; -- 查看表结构
PDU> unload tab customers; -- 导出单表为 CSV📖 完整指南: pduzc.com/quickstart
问题:PostgreSQL 因系统目录损坏无法启动。
解决方案:
PDU> b; -- 初始化元数据
PDU> use mydb;
PDU> set public;
PDU> unload tab customers; -- 导出单表
PDU> unload sch public; -- 导出模式中的所有表
PDU> unload ddl; -- 生成当前模式的 DDL
PDU> unload copy; -- 为已导出的 CSV 生成 COPY 脚本结果:CSV 文件导出到当前目录,即使 PostgreSQL 无法启动。
问题:执行了关键的 DELETE 语句,需要找回已删除的行。
解决方案:
PDU> use production;
PDU> set public;
PDU> param restype delete; -- 扫描 DELETE 记录
PDU> param resmode tx; -- 按事务号恢复
PDU> scan orders; -- 扫描 WAL 查找已删除的行
PDU> restore del <TxID>; -- 根据扫描结果恢复指定事务结果:从 WAL 归档中恢复已删除的行并导出为 CSV。
如需按时间区间恢复,先执行 param resmode time;,再设置 param starttime <YYYY-MM-DD HH:MM:SS>; 和 param endtime <YYYY-MM-DD HH:MM:SS>;,之后执行 scan orders; 与 restore del all;。
问题:一个 UPDATE 语句用错误的值修改了数百行。
解决方案:
PDU> use production;
PDU> set public;
PDU> param restype update; -- 扫描 UPDATE 记录
PDU> param resmode tx; -- 按事务号恢复
PDU> scan users;
PDU> restore upd <TxID>; -- 恢复 UPDATE 前的值结果:从 WAL 恢复 UPDATE 语句执行前的原始值。
📖 更多场景: pduzc.com/docs/instant-recovery
| 命令 | 说明 |
|---|---|
| b; | 从 PGDATA 引导元数据 |
| use <数据库>; | 切换到指定数据库 |
| set <模式>; | 切换到指定模式 |
| \l; | 列出所有数据库 |
| \dn; | 列出当前数据库的模式 |
| \dt; | 列出当前模式的表 |
| \d+ <表>; | 查看表结构(列、类型) |
| \d <表>; | 查看表列类型 |
| u tab <表>; / unload tab <表>; | 导出单表为 CSV |
| u sch <模式>; / unload sch <模式>; | 导出指定模式的所有表 |
| u ddl; / unload ddl; | 生成当前模式的 DDL |
| u copy; / unload copy; | 为已导出的 CSV 生成 COPY 语句 |
| scan <表>; | 扫描单表的 DELETE/UPDATE WAL 记录 |
| scan manual; | 从 manual 目录初始化元数据 |
| scan drop; | 扫描被 DROP/TRUNCATE 的表结构 |
| meta tab <表>; | 将指定表结构写入 tab.config |
| meta sch <模式>; | 将指定模式下的所有表结构写入 tab.config |
| restore del <TxID>; | 恢复扫描结果中的指定 DELETE 事务 |
| restore upd <TxID>; | 恢复扫描结果中的指定 UPDATE 事务 |
| restore del all; / restore upd all; | 在时间区间模式下恢复扫描结果 |
| add <filenode> <表名> <字段类型列表>; | 手动添加表信息;数据文件需放入 restore/datafile |
| restore db <库名> <路径>; | 初始化自定义数据库目录(Pro/Enterprise) |
| dropscan idx; / ds idx; | 获取被 DROP 数据页索引(Pro/Enterprise) |
| dropscan; / ds; | 按 tab.config 进行磁盘扫描恢复(Pro/Enterprise) |
| dropscan iso; / ds iso; | 按 tab.config 从镜像文件恢复(Pro/Enterprise) |
| dropscan repair; / ds repair; | 修复此前扫描失败的 TOAST 表恢复(Pro/Enterprise) |
| dropscan clean; / ds clean; | 清理 restore/dropscan 目录(Pro/Enterprise) |
| dropscan copy; / ds copy; | 为 DropScan 输出生成 COPY 命令(Pro/Enterprise) |
| p startwal <WAL>; / param startwal <WAL>; | 设置 WAL 扫描起始文件 |
| p endwal <WAL>; / param endwal <WAL>; | 设置 WAL 扫描结束文件 |
| p starttime <YYYY-MM-DD HH:MM:SS>; / param starttime <YYYY-MM-DD HH:MM:SS>; | 设置时间区间恢复起始时间 |
| p endtime <YYYY-MM-DD HH:MM:SS>; / param endtime <YYYY-MM-DD HH:MM:SS>; | 设置时间区间恢复结束时间 |
| p resmode tx|time; / param resmode tx|time; | 设置恢复模式 |
| p restype delete|update; / param restype delete|update; | 设置恢复类型 |
| p exmode csv|sql; / param exmode csv|sql; | 设置导出格式 |
| p encoding utf8|gbk; / param encoding utf8|gbk; | 设置输出编码 |
| p isomode on|off; / param isomode on|off; | 设置镜像保存模式 |
| reset <参数名>; / reset all; | 重置指定参数或全部参数 |
| show; | 查看当前参数 |
| t; | 查看支持的数据类型 |
| exit; / \q; | 退出 PDU |
所有交互式命令都必须以 ; 结尾。
📖 完整参考: pduzc.com/docs
| 参数 | 说明 | 示例 |
|---|---|---|
| PGDATA | PostgreSQL 数据目录路径 | /var/lib/postgresql/15/main |
| ARCHIVE_DEST | WAL 归档目录(用于 DELETE/UPDATE 恢复) | /var/lib/postgresql/wal_archive |
📚 完整指南请访问 pduzc.com/docs/instant-recovery
| 场景 | 说明 | 官方指南 |
|---|---|---|
| 实例损坏 | PostgreSQL 无法启动——直接从磁盘提取数据 | 打开指南 → |
| 数据文件损坏 | 特定数据文件损坏——抢救可访问的数据 | 打开指南 → |
| 数据库损坏 | 整个数据库损坏——重建目录和数据 | 打开指南 → |
| 记录被删除 | 行被误删——重放 WAL 以恢复 | 打开指南 → |
| 记录被更新 | 执行了错误的 UPDATE——用 WAL 分析回滚 | 打开指南 → |
| 无备份:DELETE / DROP TABLE / DROP DATABASE | 对比 PITR、pg_dirtyread、pg_filedump、PDU 与专业恢复服务 | 查看决策指南 → |
问:make 失败,提示 "lz4.h: No such file or directory"
答:安装 LZ4 开发库:
# Ubuntu/Debian
sudo apt-get install liblz4-dev zlib1g-dev
# RHEL/CentOS
sudo yum install lz4-devel zlib-devel问:如何更改 PostgreSQL 版本目标?
答:编辑 basic.h 并修改 PG_VERSION_NUM:
sed -i 's/#define PG_VERSION_NUM [0-9]\+/#define PG_VERSION_NUM 16/g' basic.h
make clean && make问:PDU 显示 "PGDATA directory not found"
答:验证 pdu.ini 中的 PGDATA 路径是否正确且可访问:
ls -la /var/lib/postgresql/15/main问:WAL 扫描(scan <表>)返回无结果
答:确保:
问:导出的 CSV 文件包含乱码
答:检查数据库编码:
PDU> \l; -- 查看编码列CSV 将使用数据库的编码。如需转换:
iconv -f UTF8 -t GBK input.csv > output.csv问:卸载大表耗时太长
答:PDU 执行全表扫描。为获得更好性能:
sudo apt-get update
sudo apt-get install build-essential liblz4-dev zlib1g-devsudo yum install gcc lz4-devel zlib-devel欢迎社区贡献!您可以通过以下方式帮助:
发现错误或有功能请求?提交问题
帮助让 PDU 惠及更多社区,翻译文档。
本项目采用 Apache License, Version 2.0 许可。完整文本见 LICENSE。
历史说明:早期草案使用 Business Source License 1.1。从本版本开始,PDU 在 Apache-2.0 下 100% 开源,以实现最大的社区兼容性。
本项目包含来自以下项目的派生代码:
完整归属见 NOTICE。
张晨 (Jay) 版权所有 © 2024-2025
如果 PDU 帮助您恢复了关键数据或拯救了您的一天,请考虑:
⭐ 在 GitHub 上给仓库加星 📢 与您的团队分享 PDU 🐛 报告问题以帮助改进工具 💬 加入讨论并分享您的使用案例
| Back | FazBrowse Home | New Git URL |