adamreeve commented on code in PR #400:
URL: https://github.com/apache/arrow-dotnet/pull/400#discussion_r3717967813


##########
src/Apache.Arrow/Arrays/StringViewArray.cs:
##########
@@ -22,33 +25,56 @@
 
 namespace Apache.Arrow
 {
-    public class StringViewArray : BinaryViewArray, IReadOnlyList<string>
+    public class StringViewArray(ArrayData data) : 
BinaryViewArray(ArrowTypeId.StringView, data), IReadOnlyList<string?>
     {
-        public static readonly Encoding DefaultEncoding = Encoding.UTF8;
+        public static Encoding DefaultEncoding { get; } = new 
UTF8Encoding(false);

Review Comment:
   I'm a little concerned that this could be a breaking behaviour change, but I 
checked that decoding is forgiving of whether or not a BOM is present, 
regardless of the `encoderShouldEmitUTF8Identifier` parameter. So I think this 
is OK.
   
   I notice that we disable testing the `StringView` type in our C Data 
Interface Python integration tests: 
https://github.com/apache/arrow-dotnet/blob/dc2c56617f5612cfd77750ba79b21f0e9a2ec5e6/test/Apache.Arrow.Tests/CDataInterfacePythonTests.cs#L772-L773
   
   Did you find that emitting the BOM was causing a problem when integrating 
with another Arrow implementation?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to