Files
machine-learning/01_getting_started/01_getting_started.ipynb
T
2024-09-25 18:29:02 +08:00

263 KiB

实验环境配置及基础编程训练

目录

  • 实验配置补充说明
  • Python 基础操作
    • Numpy
    • Matplotlib
  • PyTorch 基础操作
    • 基础数据操作
  • 数据预处理
  • 查阅文档
  • 线性代数

实验配置补充说明

命令行

基础安装配置已在另一单独文档讲解,如果能正常运行到这里,说明实验环境已经基本配置成功了。

即使之后的实验提示缺乏必要的软件包,也可以在 JupyterLab 集成的命令行终端中补充安装。

注意:任何在 JupyterLab 集成的命令行终端中补充安装的软件包都会被安装到与启动 JupyterLab 相同的环境中。

具体操作步骤:

  1. File -> New -> Terminal;或 File -> New Launcher,然后选择 Terminal
  2. 正常使用命令行工具更改运行环境

笔记中调用命令

在Jupyter笔记本文件中也是可以直接调用终端命令的,例如查看当前环境的Python版本:

In [1]:
!python --version
Python 3.8.16

调试器

JupyterLab也集成了调试器:使用前需要先点击右上的Enable Debugger按钮。

然后可以正常使用设置断点、单步执行、查看变量等操作。

Python 基础操作

Python本身是一种强大的通用编程语言,但在一些流行的库(numpy、scipy、matplotlib)的帮助下,它成为一种更加强大的科学计算环境。

我们希望你们中的许多人对Python和numpy有一些经验;对于其余的人,本节将作为Python编程语言和Python在科学计算中的应用的快速入门课程。

Python是一种高级的、动态类型的多范式编程语言。人们常说Python代码几乎就是伪代码,因为它允许你用很少的几行代码来表达非常强大的思想,同时又非常可读。

基础数据类型

数值型

整数和浮点数的工作方式与其他语言几乎一样。

In [2]:
x = 3
print(x, type(x))
3 <class 'int'>
In [3]:
print(x + 1)   # Addition
print(x - 1)   # Subtraction
print(x * 2)   # Multiplication
print(x ** 2)  # Exponentiation
4
2
6
9
In [4]:
x += 1
print(x)
x *= 2
print(x)
4
8
In [5]:
y = 2.5
print(type(y))
print(y, y + 1, y * 2, y ** 2)
<class 'float'>
2.5 3.5 5.0 6.25

注意,与许多语言不同,Python 没有单数增量 (x++) 或减量 (x--) 操作符。

Python 也有长整数和复数的内置类型;如有疑问,可查看文档细节。

布尔型

Python 实现了布尔逻辑的所有常用运算符,但使用英文单词而不是符号 (&&, ||, 等等)。

In [6]:
t, f = True, False
print(type(t))
<class 'bool'>

现在我们来看看这些操作:

In [7]:
print(t and f) # Logical AND;
print(t or f)  # Logical OR;
print(not t)   # Logical NOT;
print(t != f)  # Logical XOR;
False
True
False
True

字符串

字符串不区分单引号或双引号:

In [8]:
hello = 'hello'   # String literals can use single quotes
world = "world"   # or double quotes; it does not matter
print(hello, len(hello))
hello 5

字符串拼接:

In [9]:
hw = hello + ' ' + world  # String concatenation
print(hw)
hello world
In [10]:
hw12 = '{} {} {}'.format(hello, world, 12)  # string formatting
print(hw12)
hello world 12

最新版Python里(3.7以上)推荐使用“f-string”格式化字符串:

In [11]:
print(f'{hello}, {world} {12}')
hello, world 12

字符串对象有很多有用的方法;例如:

In [12]:
s = "hello"
print(s.capitalize())  # Capitalize a string
print(s.upper())       # Convert a string to uppercase; prints "HELLO"
print(s.rjust(7))      # Right-justify a string, padding with spaces
print(s.center(7))     # Center a string, padding with spaces
print(s.replace('l', '(ell)'))  # Replace all instances of one substring with another
print('  world '.strip())  # Strip leading and trailing whitespace
Hello
HELLO
  hello
 hello 
he(ell)(ell)o
world

你可以在文档中找到所有字符串方法的列表。

容器

Python 包括几种内置的容器类型:列表、字典、集合和图元。

列表

列表相当于数组,但是可以调整大小,并且可以包含不同类型的元素。

In [13]:
xs = [3, 1, 2]   # Create a list
print(xs, xs[2])
print(xs[-1])     # Negative indices count from the end of the list; prints "2"
[3, 1, 2] 2
2
In [14]:
xs[2] = 'foo'    # Lists can contain elements of different types
print(xs)
[3, 1, 'foo']
In [15]:
xs.append('bar') # Add a new element to the end of the list
print(xs)  
[3, 1, 'foo', 'bar']
In [16]:
x = xs.pop()     # Remove and return the last element of the list
print(x, xs)
bar [3, 1, 'foo']

像往常一样,你可以在文档中找到关于列表的所有细节。

数据切片

除了每次访问列表元素之外,Python 还提供了简洁的语法来访问子列表;这被称为切片。

In [17]:
nums = list(range(5))    # range is a built-in function that creates a list of integers
print(nums)         # Prints "[0, 1, 2, 3, 4]"
print(nums[2:4])    # Get a slice from index 2 to 4 (exclusive); prints "[2, 3]"
print(nums[2:])     # Get a slice from index 2 to the end; prints "[2, 3, 4]"
print(nums[:2])     # Get a slice from the start to index 2 (exclusive); prints "[0, 1]"
print(nums[:])      # Get a slice of the whole list; prints ["0, 1, 2, 3, 4]"
print(nums[:-1])    # Slice indices can be negative; prints ["0, 1, 2, 3]"
nums[2:4] = [8, 9] # Assign a new sublist to a slice
print(nums)         # Prints "[0, 1, 8, 9, 4]"
[0, 1, 2, 3, 4]
[2, 3]
[2, 3, 4]
[0, 1]
[0, 1, 2, 3, 4]
[0, 1, 2, 3]
[0, 1, 8, 9, 4]

遍历

遍历列表元素在Python中也非常简洁:

In [18]:
animals = ['cat', 'dog', 'monkey']
for animal in animals:
    print(animal)
cat
dog
monkey

如果你想访问一个循环体内每个元素的索引,请使用内置的enumerate函数。

In [19]:
animals = ['cat', 'dog', 'monkey']
for idx, animal in enumerate(animals):
    print('#{}: {}'.format(idx + 1, animal))
#1: cat
#2: dog
#3: monkey

列表理解

在编程时,我们经常想把一种类型的数据转换成另一种类型的数据。作为一个简单的例子,考虑以下计算平方数的代码。

In [20]:
nums = [0, 1, 2, 3, 4]
squares = []
for x in nums:
    squares.append(x ** 2)
print(squares)
[0, 1, 4, 9, 16]

你可以用列表理解法使这段代码更简单。

In [21]:
nums = [0, 1, 2, 3, 4]
squares = [x ** 2 for x in nums]
print(squares)
[0, 1, 4, 9, 16]

列表理解也可以包含条件:

In [22]:
nums = [0, 1, 2, 3, 4]
even_squares = [x ** 2 for x in nums if x % 2 == 0]
print(even_squares)
[0, 4, 16]

字典

dictionary 存储对 (key, value),类似于 Java 中的 Map 或 Javascript 中的对象。

In [23]:
d = {'cat': 'cute', 'dog': 'furry'}  # Create a new dictionary with some data
print(d['cat'])       # Get an entry from a dictionary; prints "cute"
print('cat' in d)     # Check if a dictionary has a given key; prints "True"
cute
True
In [24]:
d['fish'] = 'wet'    # Set an entry in a dictionary
print(d['fish'])      # Prints "wet"
wet
In [31]:
import traceback
try:
    print(d['monkey'])  # KeyError: 'monkey' not a key of d
except Exception as e:
    traceback.print_exc()
Traceback (most recent call last):
  File "/var/folders/zv/bzxgq5_j5f9gkpm84rg80kph0000gn/T/ipykernel_31837/1743455973.py", line 3, in <module>
    print(d['monkey'])  # KeyError: 'monkey' not a key of d
KeyError: 'monkey'
In [32]:
print(d.get('monkey', 'N/A'))  # Get an element with a default; prints "N/A"
print(d.get('fish', 'N/A'))    # Get an element with a default; prints "wet"
N/A
wet
In [33]:
del d['fish']        # Remove an element from a dictionary
print(d.get('fish', 'N/A')) # "fish" is no longer a key; prints "N/A"
N/A

遍历字典中的关键字:

In [34]:
d = {'person': 2, 'cat': 4, 'spider': 8}
for animal, legs in d.items():
    print('A {} has {} legs'.format(animal, legs))
A person has 2 legs
A cat has 4 legs
A spider has 8 legs

字典理解与列表理解类似,但允许你轻松构建字典:

In [35]:
nums = [0, 1, 2, 3, 4]
even_num_to_square = {x: x ** 2 for x in nums if x % 2 == 0}
print(even_num_to_square)
{0: 0, 2: 4, 4: 16}

你可以在文档中找到所有你需要知道的关于字典的信息。

集合

集合是一个由不同元素组成的无序集合。

In [36]:
animals = {'cat', 'dog'}
print('cat' in animals)   # Check if an element is in a set; prints "True"
print('fish' in animals)  # prints "False"
True
False
In [37]:
animals.add('fish')      # Add an element to a set
print('fish' in animals)
print(len(animals))       # Number of elements in a set;
True
3
In [38]:
animals.add('cat')       # Adding an element that is already in the set does nothing
print(len(animals))       
animals.remove('cat')    # Remove an element from a set
print(len(animals))       
3
2

循环:对集合进行迭代的语法与对列表进行迭代的语法相同;但是由于集合是无序的,你不能对访问集合中的元素的顺序做出假设。

In [39]:
animals = {'cat', 'dog', 'fish'}
for idx, animal in enumerate(animals):
    print('#{}: {}'.format(idx + 1, animal))
#1: fish
#2: dog
#3: cat

集合理解:像列表和字典一样,我们可以用集合理解法轻松地构建集合。

In [40]:
from math import sqrt
print({int(sqrt(x)) for x in range(30)})
{0, 1, 2, 3, 4, 5}

元组

元组是一个(不可变的)有序的数值列表。元组在许多方面与列表相似;最重要的区别之一是,元组可以作为字典的键和集合的元素,而列表则不能。

In [41]:
d = {(x, x + 1): x for x in range(10)}  # Create a dictionary with tuple keys
t = (5, 6)       # Create a tuple
print(type(t))
print(d[t])       
print(d[(1, 2)])
<class 'tuple'>
5
1
In [43]:
import traceback
try:
    t[0] = 1 # TypeError: 'tuple' object does not support item assignment
except Exception as e:
    traceback.print_exc()
Traceback (most recent call last):
  File "/var/folders/zv/bzxgq5_j5f9gkpm84rg80kph0000gn/T/ipykernel_31837/743935956.py", line 3, in <module>
    t[0] = 1 # TypeError: 'tuple' object does not support item assignment
TypeError: 'tuple' object does not support item assignment

函数

Python 函数是用 def 关键字定义的。

In [44]:
def sign(x):
    if x > 0:
        return 'positive'
    elif x < 0:
        return 'negative'
    else:
        return 'zero'

for x in [-1, 0, 1]:
    print(sign(x))
negative
zero
positive

我们可以像这样定义函数来接受可选的关键字参数。

In [45]:
def hello(name, loud=False):
    if loud:
        print('HELLO, {}'.format(name.upper()))
    else:
        print('Hello, {}!'.format(name))

hello('Bob')
hello('Fred', loud=True)
Hello, Bob!
HELLO, FRED

类

在Python中定义类的语法与其他面向对象语言类似。

In [46]:
class Greeter:

    # Constructor
    def __init__(self, name):
        self.name = name  # Create an instance variable

    # Instance method
    def greet(self, loud=False):
        if loud:
          print('HELLO, {}'.format(self.name.upper()))
        else:
          print('Hello, {}!'.format(self.name))

g = Greeter('Fred')  # Construct an instance of the Greeter class
g.greet()            # Call an instance method; prints "Hello, Fred"
g.greet(loud=True)   # Call an instance method; prints "HELLO, FRED!"
Hello, Fred!
HELLO, FRED

Numpy

Numpy是Python中科学计算的核心库。它提供了一个高性能的多维数组对象,以及处理这些数组的工具。

要使用Numpy,我们首先需要导入numpy包。

In [47]:
import numpy as np

数组

numpy数组是一个由数值组成的网格,所有的数值都是相同的类型,并由一个非负整数的元组来索引。维数的数量是数组的等级;数组的形状是一个整数的元组,给出数组在每个维度上的大小。

我们可以从嵌套的Python列表中初始化numpy数组,并使用方括号访问元素。

In [48]:
a = np.array([1, 2, 3])  # Create a rank 1 array
print(type(a), a.shape, a[0], a[1], a[2])
a[0] = 5                 # Change an element of the array
print(a)                  
<class 'numpy.ndarray'> (3,) 1 2 3
[5 2 3]
In [49]:
b = np.array([[1,2,3],[4,5,6]])   # Create a rank 2 array
print(b)
[[1 2 3]
 [4 5 6]]
In [50]:
print(b.shape)
print(b[0, 0], b[0, 1], b[1, 0])
(2, 3)
1 2 4

Numpy还提供了许多创建数组的函数。

In [51]:
a = np.zeros((2,2))  # Create an array of all zeros
print(a)
[[0. 0.]
 [0. 0.]]
In [52]:
b = np.ones((1,2))   # Create an array of all ones
print(b)
[[1. 1.]]
In [53]:
c = np.full((2,2), 7) # Create a constant array
print(c)
[[7 7]
 [7 7]]
In [54]:
d = np.eye(2)        # Create a 2x2 identity matrix
print(d)
[[1. 0.]
 [0. 1.]]
In [55]:
e = np.random.random((2,2)) # Create an array filled with random values
print(e)
[[0.00480263 0.8985954 ]
 [0.19113444 0.57867786]]

数组索引

切片:与Python列表类似,numpy数组可以被切片。由于数组可能是多维的,你必须为数组的每个维度指定一个分片。

In [56]:
import numpy as np

# Create the following rank 2 array with shape (3, 4)
# [[ 1  2  3  4]
#  [ 5  6  7  8]
#  [ 9 10 11 12]]
a = np.array([[1,2,3,4], [5,6,7,8], [9,10,11,12]])

# Use slicing to pull out the subarray consisting of the first 2 rows
# and columns 1 and 2; b is the following array of shape (2, 2):
# [[2 3]
#  [6 7]]
b = a[:2, 1:3]
print(b)
[[2 3]
 [6 7]]

一个数组的片断是对同一数据的视图,所以修改它将修改原始数组。

In [57]:
print(a[0, 1])
b[0, 0] = 77    # b[0, 0] is the same piece of data as a[0, 1]
print(a[0, 1]) 
2
77

你也可以把整数索引和片断索引混合起来。但是,这样做会产生一个比原数组秩更低的数组。

In [58]:
# Create the following rank 2 array with shape (3, 4)
a = np.array([[1,2,3,4], [5,6,7,8], [9,10,11,12]])
print(a)
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]

访问数组中间行的数据的两种方法。将整数索引与分片混合使用会产生一个较低秩的数组,而只使用分片会产生一个与原数组秩相同的数组。

In [59]:
row_r1 = a[1, :]    # Rank 1 view of the second row of a  
row_r2 = a[1:2, :]  # Rank 2 view of the second row of a
row_r3 = a[[1], :]  # Rank 2 view of the second row of a
print(row_r1, row_r1.shape)
print(row_r2, row_r2.shape)
print(row_r3, row_r3.shape)
[5 6 7 8] (4,)
[[5 6 7 8]] (1, 4)
[[5 6 7 8]] (1, 4)
Warning:
Output truncated. This notebook contains too many cells to display efficiently.