| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Signed-off-by: ebolblga <kkochanovskiy@gmail.com>
| Back | FazBrowse Home | New Git URL |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Choose a reason Spam Abuse Off Topic Outdated Duplicate Resolved Low Qualitywhy remove the explicit schema declaration?
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Choose a reason Spam Abuse Off Topic Outdated Duplicate Resolved Low QualityLet's say I have Pandas DataFrame like this:
If I output column dtypes then feature_x would still be object (array(int)) and feature_y would be converted to float by underlying NumPy because of the Null.
By default, pandas uses NumPy data types, which do not support missing values in integer arrays. If you create a Series or DataFrame column with integers and include a null value (e.g., None or np.nan), pandas will upcast the column to a floating-point type (float64) to accommodate the missing value.
If I later pass this dataframe WITH forced schema, PyArrow will see that I want to cast float to array(int) and throw an error. If I don't pass the schema though, it will infer type itself and work as expected.
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.