Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

Question

I'm new to NLTK. I'm getting this error and I've searched around for encoding/decoding and specifically the UnicodeDecodeError but this error seems specific to the NLTK source code.

Here's the error:

Traceback (most recent call last):
  File "A:\Python\Projects\Test\main.py", line 2, in <module>
    print(pos_tag(word_tokenize("John's big idea isn't all that bad.")))
  File "A:\Python\Python\lib\site-packages\nltk\tag\__init__.py", line 100, in pos_tag
    tagger = load(_POS_TAGGER)
  File "A:\Python\Python\lib\site-packages\nltk\data.py", line 779, in load
    resource_val = pickle.load(opened_resource)
UnicodeDecodeError: 'ascii' codec can't decode byte 0xcb in position 0: ordinal not in range(128)

How do I go around fixing this error?

Here's what causes the error:

from nltk import pos_tag, word_tokenize
print(pos_tag(word_tokenize("John's big idea isn't all that bad.")))

There's nothing in the code you show here that would generate the error. Print the repr of the string you're passing. — Mark Ransom
– Mark Ransom, Commented Aug 25, 2014 at 21:08
@MarkRansom I don't know what you mean, the function pos_tag is causing the error. I think the encoding error is generated on the pickle.load function. I'm not sure what to do. — user3422952
– user3422952, Commented Aug 25, 2014 at 21:22

Simone Dagli Orti · Accepted Answer · 2015-01-14 17:06:43Z

5

try this... NLTK 3.0.1 with Python 2.7.x

import io
f = io.open(txtFile, 'rU', encoding='utf-8')

answered Jan 14, 2015 at 17:06

Simone Dagli Orti

4766 silver badges5 bronze badges

Sign up to request clarification or add additional context in comments.

2 Comments

Axol Otl Over a year ago

Worked like a charm! I use nltk 3.1 and Python 2.7.x.

minocha Over a year ago

This is great ! can you also explain why using io solves the problem ?

LuckyMatina · Accepted Answer · 2015-03-04 19:08:53Z

4

I had the same problem with you. I use Python 3.4 in Windows 7.

I had installed the "nltk-3.0.0.win32.exe" (from here). But when i installed the "nltk-3.0a4.win32.exe" (from here), my problem with nltk.pos_tag was solved. Check it.

EDIT: If the second link doesn't work, you can look here.

edited Mar 4, 2015 at 19:08

answered Sep 28, 2014 at 17:39

LuckyMatina

416 bronze badges

1 Comment

Arnab Chakraborty Over a year ago

the second link seems to be broken. Do you have any alternate links?

Community · Accepted Answer · 2017-05-23 12:00:50Z

-2

Duplicate: NLTK 3 POS_TAG throws UnicodeDecodeError

Long story short: NLTK isn't compatible with Python 3. You have to use NLTK 3 which sounds a bit experimental at this point.

edited May 23, 2017 at 12:00

CommunityBot

11 silver badge

answered Sep 3, 2014 at 2:22

Dave

1

1 Comment

Gustavo Puma Over a year ago

I am using NLTK 3 and Python 3.4 and still get this error.

Shivamshaz · Accepted Answer · 2014-09-05 09:36:40Z

-2

Try using the module "textclean"

>>> pip install textclean

Python code

from textclean.textclean import textclean
text = textclean.clean("John's big idea isn't all that bad.")
print pos_tag(word_tokenize(text))

answered Sep 5, 2014 at 9:36

Shivamshaz

2802 gold badges3 silver badges10 bronze badges

1 Comment

Eevee Over a year ago

this module sounds like a horrible idea. it's particularly bad here, because the error is occurring when trying to decode a pickle — a structured data format that you will irreparably destroy if you try to blindly "clean" it.

Collectives™ on Stack Overflow

Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

4 Answers 4

2 Comments

1 Comment

1 Comment

1 Comment

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

4 Answers 4

2 Comments

1 Comment

1 Comment

1 Comment

Your Answer

Sign up or log in

Post as a guest

Linked

Related