joseluisll opened a new pull request, #8789:
URL: https://github.com/apache/hadoop/pull/8789

   ### Description of PR
   
   [HDFS-17986](https://issues.apache.org/jira/browse/HDFS-17986): with 
short-circuit reads enabled and checksum verification disabled (as HBase does 
by default), a replica with a corrupt meta file header is never reported. The 
client fails to parse the header and falls back to a remote read without 
checksums, so the DataNode never opens the meta file.
   
   `DataNode#requestShortCircuitFdsForRead` now reads the meta header before 
passing the file descriptors. If the read fails on a finalized replica, the 
DataNode calls `handleBadBlock`. The descriptors are still returned, because 
failing the request would make the client disable short-circuit reads for the 
whole DataNode.
   
   This supersedes the approach in HDFS-17179 / #8310.
   
   ### How was this patch tested?
   
   **New unit test `TestShortCircuitCorruptMetaHeader`** (needs native domain 
sockets). Each case:
   1. Starts a MiniDFSCluster with short-circuit reads enabled and writes a 
file.
   2. Corrupts the meta header of one replica, either with an invalid checksum 
type or by truncating the meta file.
   3. Reads the file from that replica, with checksum verification on or off.
   4. Checks that the read behaves as expected and that the NameNode marks the 
replica corrupt.
   
   It runs with replication 1, 3, 5 and 7 (16 cases). Without the fix, the 8 
cases with checksums off fail: the read succeeds but the replica is never 
reported. With the fix, all 16 pass.
   
   **Existing tests:** the short-circuit and corrupt-metadata suites pass 
locally (`TestShortCircuitLocalRead`, `TestBlockReaderFactory`, 
`TestShortCircuitCache`, `TestBlockReaderLocal`, `TestCorruptMetadataFile`, 
`TestScrLazyPersistFiles`).
   
   **End-to-end, before the fix:** HBase 2.6.7 on HDFS 3.5.0 with default 
settings. After corrupting the header of one block of a large HFile, HBase read 
all rows without errors on every scan, while `hdfs fsck` kept reporting the 
file as HEALTHY. Details are in the JIRA.
   
   ### For code changes:
   
   - [x] Does the title of this PR start with the corresponding JIRA issue id 
(e.g. 'HADOOP-17799. Your PR title ...')?
   - [ ] Object storage: N/A
   - [x] New dependencies: none
   - [x] `LICENSE`, `LICENSE-binary`, `NOTICE-binary`: N/A
   
   ### AI Tooling
   
   - [x] The PR includes the phrase "Contains content generated by Claude Code"
   - [x] My use of AI contributions follows the ASF legal policy 
https://www.apache.org/legal/generative-tooling.html
   
   Contains content generated by Claude Code.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to