joseluisll opened a new pull request, #8789: URL: https://github.com/apache/hadoop/pull/8789
### Description of PR [HDFS-17986](https://issues.apache.org/jira/browse/HDFS-17986): with short-circuit reads enabled and checksum verification disabled (as HBase does by default), a replica with a corrupt meta file header is never reported. The client fails to parse the header and falls back to a remote read without checksums, so the DataNode never opens the meta file. `DataNode#requestShortCircuitFdsForRead` now reads the meta header before passing the file descriptors. If the read fails on a finalized replica, the DataNode calls `handleBadBlock`. The descriptors are still returned, because failing the request would make the client disable short-circuit reads for the whole DataNode. This supersedes the approach in HDFS-17179 / #8310. ### How was this patch tested? **New unit test `TestShortCircuitCorruptMetaHeader`** (needs native domain sockets). Each case: 1. Starts a MiniDFSCluster with short-circuit reads enabled and writes a file. 2. Corrupts the meta header of one replica, either with an invalid checksum type or by truncating the meta file. 3. Reads the file from that replica, with checksum verification on or off. 4. Checks that the read behaves as expected and that the NameNode marks the replica corrupt. It runs with replication 1, 3, 5 and 7 (16 cases). Without the fix, the 8 cases with checksums off fail: the read succeeds but the replica is never reported. With the fix, all 16 pass. **Existing tests:** the short-circuit and corrupt-metadata suites pass locally (`TestShortCircuitLocalRead`, `TestBlockReaderFactory`, `TestShortCircuitCache`, `TestBlockReaderLocal`, `TestCorruptMetadataFile`, `TestScrLazyPersistFiles`). **End-to-end, before the fix:** HBase 2.6.7 on HDFS 3.5.0 with default settings. After corrupting the header of one block of a large HFile, HBase read all rows without errors on every scan, while `hdfs fsck` kept reporting the file as HEALTHY. Details are in the JIRA. ### For code changes: - [x] Does the title of this PR start with the corresponding JIRA issue id (e.g. 'HADOOP-17799. Your PR title ...')? - [ ] Object storage: N/A - [x] New dependencies: none - [x] `LICENSE`, `LICENSE-binary`, `NOTICE-binary`: N/A ### AI Tooling - [x] The PR includes the phrase "Contains content generated by Claude Code" - [x] My use of AI contributions follows the ASF legal policy https://www.apache.org/legal/generative-tooling.html Contains content generated by Claude Code. 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
