FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Merge pull request #585 from Kaggle/category_encoders_from_git · feibyte/docker-python@6a11b97 · GitHub

Commit 6a11b97

Browse files
authored
Merge pull request Kaggle#585 from Kaggle/category_encoders_from_git
Install category_encoders from GitHub
2 parents 1427897 + edfc261 commit 6a11b97

2 files changed

Lines changed: 18 additions & 1 deletion

File tree

‎Dockerfile‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -346,7 +346,8 @@ RUN pip install --upgrade cython && \
346346
pip install fasttext && \
347347
apt-get install -y libhunspell-dev && pip install hunspell && \
348348
pip install annoy && \
349-
pip install category_encoders && \
349+
# Need to use CountEncoder from category_encoders before it's officially released
350+
pip install git+https://github.com/scikit-learn-contrib/categorical-encoding.git && \
350351
# Newer version crashes (latest = 1.14.0) when running tensorflow.
351352
# python -c "from google.cloud import bigquery; import tensorflow". This flow is common because bigquery is imported in kaggle_gcp.py
352353
# which is loaded at startup.

‎tests/test_category_encoders.py‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,16 @@
1+
import unittest
2+
3+
## Need to make sure we have CountEncoder available from the category_encoders library
4+
class TestCategoryEncoders(unittest.TestCase):
5+
def test_count_encoder(self):
6+
7+
from category_encoders import CountEncoder
8+
import pandas as pd
9+
10+
encoder = CountEncoder(cols="data")
11+
12+
data = pd.DataFrame([1, 2, 3, 1, 4, 5, 3, 1], columns=["data"])
13+
14+
encoded = encoder.fit_transform(data)
15+
self.assertTrue((encoded.data == [3, 1, 2, 3, 1, 1, 2, 3]).all())
16+

0 commit comments

Comments
 (0)

Back | FazBrowse Home | New Git URL