1.9 KiB
id, source_exercise_id, title, section, source_path, source_repo, source_commit, student_visible_solution, has_private_solution, skip
| id | source_exercise_id | title | section | source_path | source_repo | source_commit | student_visible_solution | has_private_solution | skip |
|---|---|---|---|---|---|---|---|---|---|
| practical-python-1.28 | 1.28 | Other kinds of 'files' | 1.6 File Management | 01_Introduction/06_Files.md | https://github.com/dabeaz-course/practical-python | 93dca856b41c61a0a0f85ae334116e4c125629ea | false | false | false |
Exercise 1.28: Other kinds of "files"
Source: Practical Python Programming,
01_Introduction/06_Files.md.
Exercise 1.28: Other kinds of "files"
What if you wanted to read a non-text file such as a gzip-compressed
datafile? The builtin open() function won’t help you here, but
Python has a library module gzip that can read gzip compressed
files.
Try it:
>>> import gzip
>>> with gzip.open('Data/portfolio.csv.gz', 'rt') as f:
for line in f:
print(line, end='')
... look at the output ...
>>>
Note: Including the file mode of 'rt' is critical here. If you forget that,
you'll get byte strings instead of normal text strings.
Commentary: Shouldn't we being using Pandas for this?
Data scientists are quick to point out that libraries like Pandas already have a function for reading CSV files. This is true--and it works pretty well. However, this is not a course on learning Pandas. Reading files is a more general problem than the specifics of CSV files. The main reason we're working with a CSV file is that it's a familiar format to most coders and it's relatively easy to work with directly--illustrating many Python features in the process. So, by all means use Pandas when you go back to work. For the rest of this course however, we're going to stick with standard Python functionality.
Contents | Previous (1.5 Lists) | Next (1.7 Functions)