This is an automated email from the ASF dual-hosted git repository.

gurwls223 pushed a commit to branch branch-3.1
in repository https://gitbox.apache.org/repos/asf/spark.git


The following commit(s) were added to refs/heads/branch-3.1 by this push:
     new 8236c5f  [SPARK-34803][PYSPARK] Pass the raised ImportError if pandas 
or pyarrow fail to import
8236c5f is described below

commit 8236c5f71cd1502cb8e54a8dc2c4801bcb319764
Author: John Ayad <johnhan...@gmail.com>
AuthorDate: Mon Mar 22 23:29:28 2021 +0900

    [SPARK-34803][PYSPARK] Pass the raised ImportError if pandas or pyarrow 
fail to import
    
    ### What changes were proposed in this pull request?
    
    Pass the raised `ImportError` on failing to import pandas/pyarrow. This 
will help the user identify whether pandas/pyarrow are indeed not in the 
environment or if they threw a different `ImportError`.
    
    ### Why are the changes needed?
    
    This can already happen in Pandas for example where it could throw an 
`ImportError` on its initialisation path if `dateutil` doesn't satisfy a 
certain version requirement 
https://github.com/pandas-dev/pandas/blob/0.24.x/pandas/compat/__init__.py#L438
    
    ### Does this PR introduce _any_ user-facing change?
    
    Yes, it will now show the root cause of the exception when pandas or arrow 
is missing during import.
    
    ### How was this patch tested?
    
    Manually tested.
    
    ```python
    from pyspark.sql.functions import pandas_udf
    spark.range(1).select(pandas_udf(lambda x: x))
    ```
    
    Before:
    
    ```
    Traceback (most recent call last):
      File "<stdin>", line 1, in <module>
      File "/...//spark/python/pyspark/sql/pandas/functions.py", line 332, in 
pandas_udf
        require_minimum_pyarrow_version()
      File "/.../spark/python/pyspark/sql/pandas/utils.py", line 53, in 
require_minimum_pyarrow_version
        raise ImportError("PyArrow >= %s must be installed; however, "
    ImportError: PyArrow >= 1.0.0 must be installed; however, it was not found.
    ```
    
    After:
    
    ```
    Traceback (most recent call last):
      File "/.../spark/python/pyspark/sql/pandas/utils.py", line 49, in 
require_minimum_pyarrow_version
        import pyarrow
    ModuleNotFoundError: No module named 'pyarrow'
    
    The above exception was the direct cause of the following exception:
    
    Traceback (most recent call last):
      File "<stdin>", line 1, in <module>
      File "/.../spark/python/pyspark/sql/pandas/functions.py", line 332, in 
pandas_udf
        require_minimum_pyarrow_version()
      File "/.../spark/python/pyspark/sql/pandas/utils.py", line 55, in 
require_minimum_pyarrow_version
        raise ImportError("PyArrow >= %s must be installed; however, "
    ImportError: PyArrow >= 1.0.0 must be installed; however, it was not found.
    ```
    
    Closes #31902 from johnhany97/jayad/spark-34803.
    
    Lead-authored-by: John Ayad <johnhan...@gmail.com>
    Co-authored-by: John H. Ayad <johnhan...@gmail.com>
    Co-authored-by: HyukjinKwon <gurwls...@apache.org>
    Signed-off-by: HyukjinKwon <gurwls...@apache.org>
    (cherry picked from commit ddfc75ec648d57f92f474d5820d03c37f20403dc)
    Signed-off-by: HyukjinKwon <gurwls...@apache.org>
---
 python/pyspark/sql/pandas/utils.py | 10 ++++++----
 1 file changed, 6 insertions(+), 4 deletions(-)

diff --git a/python/pyspark/sql/pandas/utils.py 
b/python/pyspark/sql/pandas/utils.py
index 9b97676..b22603a 100644
--- a/python/pyspark/sql/pandas/utils.py
+++ b/python/pyspark/sql/pandas/utils.py
@@ -26,11 +26,12 @@ def require_minimum_pandas_version():
     try:
         import pandas
         have_pandas = True
-    except ImportError:
+    except ImportError as error:
         have_pandas = False
+        raised_error = error
     if not have_pandas:
         raise ImportError("Pandas >= %s must be installed; however, "
-                          "it was not found." % minimum_pandas_version)
+                          "it was not found." % minimum_pandas_version) from 
raised_error
     if LooseVersion(pandas.__version__) < LooseVersion(minimum_pandas_version):
         raise ImportError("Pandas >= %s must be installed; however, "
                           "your version was %s." % (minimum_pandas_version, 
pandas.__version__))
@@ -47,11 +48,12 @@ def require_minimum_pyarrow_version():
     try:
         import pyarrow
         have_arrow = True
-    except ImportError:
+    except ImportError as error:
         have_arrow = False
+        raised_error = error
     if not have_arrow:
         raise ImportError("PyArrow >= %s must be installed; however, "
-                          "it was not found." % minimum_pyarrow_version)
+                          "it was not found." % minimum_pyarrow_version) from 
raised_error
     if LooseVersion(pyarrow.__version__) < 
LooseVersion(minimum_pyarrow_version):
         raise ImportError("PyArrow >= %s must be installed; however, "
                           "your version was %s." % (minimum_pyarrow_version, 
pyarrow.__version__))

---------------------------------------------------------------------
To unsubscribe, e-mail: commits-unsubscr...@spark.apache.org
For additional commands, e-mail: commits-h...@spark.apache.org

Reply via email to