[
https://issues.apache.org/jira/browse/HADOOP-19971?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jose Luis López updated HADOOP-19971:
-------------------------------------
Description:
h3. Why
Hadoop will eventually move from the javax.servlet namespace to
jakarta.servlet. When it does, every class that mentions a servlet type has to
be revisited, and every class that mentions a Jetty type has to be revisited
again when the Jetty API changes underneath it.
Several of those classes have no real connection to either. A component that
generates signing secrets was handed the web server's context even though it
only wanted somewhere to store an object. A certificate parser reported a bad
certificate as a web server error. A Netty-based shuffle handler borrowed two
header names from Jetty, and the YARN services client borrowed Jetty's URL
encoder.
This change removes those references, so those classes drop out of the later
migration entirely.
h3. What changes for anyone using Hadoop
Nothing is removed, no signature changes shape, and no dependency, version or
setting changes. Code built against today's release keeps compiling and running
without being rebuilt.
* Two entry points get a servlet-free alternative. The old ones stay, marked
deprecated, until the namespace change:
** {{SignerSecretProvider.init(Properties, ServletContext, long)}} ->
{{initialize(Properties, SecretProviderContext, long)}}
** {{CertificateUtil.parseRSAPublicKey(String)}} -> {{toRSAPublicKey(String)}}
* Providers that override {{init}} are still called directly. A provider that
extends Hadoop's rolling secret provider and calls {{super.init}} still gets
its rollover scheduler started. {{TestSignerSecretProviderCompatibility}} pins
this contract.
* Exceptions are unchanged: same type, message and cause chain from
{{parseRSAPublicKey}} and from {{JWTRedirectAuthenticationHandler}} when the
PEM is corrupt.
* The shuffle response is unchanged, byte for byte: the same {{Connection}} and
{{Keep-Alive}} header names, capitalised as before.
* The YARN services client now encodes {{user.name}} with
{{java.net.URLEncoder}}. That differs from Jetty's encoder for two characters
only: {{*}} is now sent as is where it used to be escaped, and {{~}} is now
escaped where it used to be sent as is. Every code point decodes to the same
value either way (checked across all of Unicode against Jetty 9.4.58), so the
server reads the same user name.
Things a downstream project could notice:
* Deprecation warnings where code overrides {{init}} or calls
{{parseRSAPublicKey}}.
* {{getDeclaredMethod("init", ...)}} on the shipped providers now throws
{{NoSuchMethodException}}, because they declare {{initialize}} instead.
{{getMethod}} still finds {{init}}.
* {{ZKSignerSecretProvider}} given a null ServletContext used to throw a
NullPointerException. It now keeps its ZooKeeper client in a store private to
itself and logs a WARN. Every caller in Hadoop passes a real context, and there
the client is still stored as the ServletContext attribute of the same name.
h3. What changes inside
|| || before || after ||
| modules using Jetty in main sources | 12 | 11 |
| main-source files using Jetty | 26 | 24 |
| hadoop-auth main classes naming a servlet | 14 | 11 |
Counted on trunk at a7c2bea723d.
The MapReduce shuffle module no longer uses Jetty at all. In hadoop-auth, the
secret providers and the authentication token no longer mention the servlet
API. They are initialised against a small two-method store,
{{SecretProviderContext}}, which holds whatever needs sharing; that is all the
one provider that used the servlet context ever did with it. One new
package-private class mentions the servlet API, and it exists only to keep the
old path working.
h3. What this does not do
No namespace change, no Jetty upgrade, no Jersey change, no EE environments.
The tree stays on Jetty 9.4, Jersey 2 and javax.servlet.
The authentication handler classes still mention the servlet API, deliberately.
They are the extension point that HBase, Hive, Spark, Ozone and Knox build
against, and changing it is exactly the kind of break this issue is meant to
avoid. That belongs with the namespace change, where downstream projects will
expect it.
h3. Ordering
This is preparation for the jakarta move. It depends on nothing, is based on
trunk, and nothing in HADOOP-19972 depends on it. The plan is to land it with
or just before HADOOP-19912, in the next major release.
When rebasing onto a later trunk:
* If HADOOP-19970 has landed, it declares jetty-http in the shuffle module's
pom only because of the import this change removes. Drop that declaration here.
* If HADOOP-19972 has landed, re-check the shuffle headers and the
{{user.name}} encoding against whatever Jetty classes it left in those two
places.
16 files changed. Works whether HADOOP-19912 ends up going through Jetty 12's
ee8 environment or straight to ee10, and commits the project to neither.
Contains content generated by Claude.
was:
WHY
===
Hadoop will eventually move from the javax.servlet namespace to
jakarta.servlet. When it does, every class that mentions a servlet type has to
be revisited, and every class that mentions a Jetty type has to be revisited
again when the Jetty API changes underneath it.
Several of those classes have no real connection to either. A component that
generates signing secrets was handed the web server's context even though it
only wanted somewhere to stash an object. A certificate parser reported a bad
certificate as a web server error. A file-transfer handler built on Netty
borrowed two header names from Jetty.
This clears those cases out. The later migration is smaller, and the classes
concerned stop being part of it at all.
WHAT CHANGES FOR ANYONE USING HADOOP
====================================
Nothing. Nothing is removed, nothing changes shape, and no dependency, version
or setting changes. Code built against today's release keeps compiling and
keeps running without being rebuilt.
Two entry points were the awkward ones. Both stay exactly as they are, marked
deprecated, with a cleaner alternative alongside. Projects that adopt the new
one are unaffected when the namespace finally moves; projects that do not are
no worse off than today. A test suite exists purely to prove the old path still
behaves - including the awkward case of a plugin that extends Hadoop's rolling
secret provider and calls up into it.
The shuffle protocol is untouched, byte for byte. Deployments see no difference.
WHAT CHANGES INSIDE
===================
before after
modules using Jetty in main sources 13 12
main-source files using Jetty 25 23
hadoop-auth main classes naming a servlet 14 11
The MapReduce shuffle module stops depending on Jetty altogether, in code and
in its build file.
In hadoop-auth, the secret providers and the authentication token no longer
mention the servlet API. They are initialised instead against a small,
two-method store that holds whatever needs sharing - which is all the one
provider that used the servlet context was ever doing with it. One new class
does mention the servlet API, and exists solely to keep the old path working.
The YARN services client and the shuffle handler each replaced a small Jetty
utility with the equivalent from the JDK or from Netty, which they already
depend on.
WHAT THIS DOES NOT DO
=====================
No namespace change, no Jetty upgrade, no Jersey change, no EE environments.
The tree stays on Jetty 9.4, Jersey 2 and javax.servlet throughout.
The authentication handler classes still mention the servlet API, deliberately.
They sit on the extension point that HBase, Hive, Spark, Ozone and Knox build
against, and changing it is precisely the kind of break this issue is designed
to avoid. It belongs with the namespace change, where downstream projects are
expecting it.
DEPENDS ON
==========
HADOOP-19970 (PR #8699), for sequencing only. That PR declares a Jetty artifact
for the shuffle module; this one removes the last use of it, and drops the
declaration in the same change.
16 files changed. Works whether HADOOP-19912 ends up going through Jetty 12's
ee8 environment or straight to ee10, and commits the project to neither.
Contains content generated by Claude.
> Remove servlet and Jetty types from classes that do not need them
> -----------------------------------------------------------------
>
> Key: HADOOP-19971
> URL: https://issues.apache.org/jira/browse/HADOOP-19971
> Project: Hadoop Common
> Issue Type: Sub-task
> Components: auth, common
> Reporter: Jose Luis López
> Assignee: Jose Luis López
> Priority: Minor
> Labels: pull-request-available
>
> h3. Why
> Hadoop will eventually move from the javax.servlet namespace to
> jakarta.servlet. When it does, every class that mentions a servlet type has
> to be revisited, and every class that mentions a Jetty type has to be
> revisited again when the Jetty API changes underneath it.
> Several of those classes have no real connection to either. A component that
> generates signing secrets was handed the web server's context even though it
> only wanted somewhere to store an object. A certificate parser reported a bad
> certificate as a web server error. A Netty-based shuffle handler borrowed two
> header names from Jetty, and the YARN services client borrowed Jetty's URL
> encoder.
> This change removes those references, so those classes drop out of the later
> migration entirely.
> h3. What changes for anyone using Hadoop
> Nothing is removed, no signature changes shape, and no dependency, version or
> setting changes. Code built against today's release keeps compiling and
> running without being rebuilt.
> * Two entry points get a servlet-free alternative. The old ones stay, marked
> deprecated, until the namespace change:
> ** {{SignerSecretProvider.init(Properties, ServletContext, long)}} ->
> {{initialize(Properties, SecretProviderContext, long)}}
> ** {{CertificateUtil.parseRSAPublicKey(String)}} -> {{toRSAPublicKey(String)}}
> * Providers that override {{init}} are still called directly. A provider that
> extends Hadoop's rolling secret provider and calls {{super.init}} still gets
> its rollover scheduler started. {{TestSignerSecretProviderCompatibility}}
> pins this contract.
> * Exceptions are unchanged: same type, message and cause chain from
> {{parseRSAPublicKey}} and from {{JWTRedirectAuthenticationHandler}} when the
> PEM is corrupt.
> * The shuffle response is unchanged, byte for byte: the same {{Connection}}
> and {{Keep-Alive}} header names, capitalised as before.
> * The YARN services client now encodes {{user.name}} with
> {{java.net.URLEncoder}}. That differs from Jetty's encoder for two characters
> only: {{*}} is now sent as is where it used to be escaped, and {{~}} is now
> escaped where it used to be sent as is. Every code point decodes to the same
> value either way (checked across all of Unicode against Jetty 9.4.58), so the
> server reads the same user name.
> Things a downstream project could notice:
> * Deprecation warnings where code overrides {{init}} or calls
> {{parseRSAPublicKey}}.
> * {{getDeclaredMethod("init", ...)}} on the shipped providers now throws
> {{NoSuchMethodException}}, because they declare {{initialize}} instead.
> {{getMethod}} still finds {{init}}.
> * {{ZKSignerSecretProvider}} given a null ServletContext used to throw a
> NullPointerException. It now keeps its ZooKeeper client in a store private to
> itself and logs a WARN. Every caller in Hadoop passes a real context, and
> there the client is still stored as the ServletContext attribute of the same
> name.
> h3. What changes inside
> || || before || after ||
> | modules using Jetty in main sources | 12 | 11 |
> | main-source files using Jetty | 26 | 24 |
> | hadoop-auth main classes naming a servlet | 14 | 11 |
> Counted on trunk at a7c2bea723d.
> The MapReduce shuffle module no longer uses Jetty at all. In hadoop-auth, the
> secret providers and the authentication token no longer mention the servlet
> API. They are initialised against a small two-method store,
> {{SecretProviderContext}}, which holds whatever needs sharing; that is all
> the one provider that used the servlet context ever did with it. One new
> package-private class mentions the servlet API, and it exists only to keep
> the old path working.
> h3. What this does not do
> No namespace change, no Jetty upgrade, no Jersey change, no EE environments.
> The tree stays on Jetty 9.4, Jersey 2 and javax.servlet.
> The authentication handler classes still mention the servlet API,
> deliberately. They are the extension point that HBase, Hive, Spark, Ozone and
> Knox build against, and changing it is exactly the kind of break this issue
> is meant to avoid. That belongs with the namespace change, where downstream
> projects will expect it.
> h3. Ordering
> This is preparation for the jakarta move. It depends on nothing, is based on
> trunk, and nothing in HADOOP-19972 depends on it. The plan is to land it with
> or just before HADOOP-19912, in the next major release.
> When rebasing onto a later trunk:
> * If HADOOP-19970 has landed, it declares jetty-http in the shuffle module's
> pom only because of the import this change removes. Drop that declaration
> here.
> * If HADOOP-19972 has landed, re-check the shuffle headers and the
> {{user.name}} encoding against whatever Jetty classes it left in those two
> places.
> 16 files changed. Works whether HADOOP-19912 ends up going through Jetty 12's
> ee8 environment or straight to ee10, and commits the project to neither.
> Contains content generated by Claude.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]