解析库之xpath，beautifulsoup，pyquery

时间 2019-12-07

标签解析 xpath beautifulsoup pyquery 繁體版

原文原文链接

xpath

1、经常使用xpath表达式

属性定位：
1. #找到class属性值为song的div标签
2. //div[@class="song"]
层级&索引定位：
1. #找到class属性值为tang的div的直系子标签ul下的第二个子标签li下的直系子标签a
2. //div[@class="tang"]/ul/li[2]/a
逻辑运算：
1. #找到href属性值为空且class属性值为du的a标签
2. //a[@href="" and @class="du"]
模糊匹配：
1. //div[contains(@class, "ng")]
2. //div[starts-with(@class, "ta")]
取文本：
1. # /表示获取某个标签下的文本内容
2. # //表示获取某个标签下的文本内容和全部子标签下的文本内容
3. //div[@class="song"]/p[1]/text()
4. //div[@class="tang"]//text()
取属性：
1. //div[@class="tang"]//li[2]/a/@href
xpath函数返回的老是一个列表

2、基础使用

下载：pip install lxml
导包：from lxml import etree
官网推荐在如今的项目中使用Beautiful Soup 4, 移植到BS4
将html文档或者xml文档转换成一个etree对象，而后调用对象中的方法查找指定的节点
本地文件：
1. tree = etree.parse('本地文件路径')
2. tree.xpath("xpath表达式")
网络数据：
1. tree = etree.HTML("网络请求到的页面数据")
2. tree.xpath("xpath表达式")
xpath插件：就能够直接将xpath表达式做用于浏览器的网页当中
安装：更多工具-》扩展程序-》开启右上角的开发者模式-》xpath插件拖动到页面便可
快捷键：启动和关闭插件 ctrl + shift + x

Beautifulsoup模块

Beautiful Soup 是一个能够从HTML或XML文件中提取数据的Python库.它可以经过你喜欢的转换器实现惯用的文档导航,查找,修改文档的方式.
Beautiful Soup会帮你节省数小时甚至数天的工做时间.你可能在寻找 Beautiful Soup3 的文档,Beautiful Soup 3 目前已经中止开发.
官网推荐在如今的项目中使用Beautiful Soup 4, 移植到BS4
官网推荐使用lxml做为解析器,由于效率更高.
在Python2.7.3以前的版本和Python3中3.2.2以前的版本,必须安装lxml或html5lib, 由于那些Python版本的标准库中内置的HTML解析方法不够稳定.
中文文档点击

1、环境安装

须要将pip源设置为国内源，阿里源、豆瓣源、网易源等
windows：
1. 打开文件资源管理器(文件夹地址栏中)
2. 地址栏上面输入 %appdata%
3. 在这里面新建一个文件夹 pip
4. 在pip文件夹里面新建一个文件叫作 pip.ini ,内容写以下便可
5. [global]
6. timeout = 6000
7. index-url = https://mirrors.aliyun.com/pypi/simple/
8. trusted-host = mirrors.aliyun.com
linux：
1. cd ~
2. mkdir ~/.pip
3. vi ~/.pip/pip.conf
4. 编辑内容，和windows如出一辙
须要安装：pip install bs4
1. bs4在使用时候须要一个第三方库，把这个库也安装一下
2. pip install lxml

2、基础使用

核心思想：将html文档转换成Beautiful对象，而后调用该对象中的属性和方法进行html文档指定内容的定位查找。javascript

一、使用流程

导包：from bs4 import BeautifulSoup
使用方式：能够将一个html文档，转化为BeautifulSoup对象，而后经过对象的方法或者属性去查找指定的节点内容
转化本地文件：soup = BeautifulSoup(open('本地文件'), 'lxml')
转化网络文件：soup = BeautifulSoup('字符串类型或者字节类型', 'lxml')
打印soup对象显示内容为html文件中的内容

二、基础巩固

根据标签名查找：
1. soup.a 只能找到第一个符合要求的标签
获取属性：
1. soup.a.attrs 获取a全部的属性和属性值，返回一个字典
2. soup.a.attrs['href'] 获取href属性
3. soup.a['href'] 也可简写为这种形式
获取内容：
1. soup.a.string
2. soup.a.text
3. soup.a.get_text()
4. 【注意】若是标签还有标签，那么string获取到的结果为None，而其它两个，能够获取文本内容
find：找到第一个符合要求的标签：
1. soup.find('a') 找到第一个符合要求的
2. soup.find('a', title="xxx")
3. soup.find('a', alt="xxx")
4. soup.find('a', class_="xxx")
5. soup.find('a', id="xxx")
find_All：找到全部符合要求的标签：
1. soup.find_All('a')
2. soup.find_All(['a','b']) 找到全部的a和b标签
3. soup.find_All('a', limit=2) 限制前两个
根据选择器选择指定的内容：
1. select:soup.select('#feng')
2. 常见的选择器：标签选择器(a)、类选择器(.)、id选择器(#)、层级选择器
3. 层级选择器：div .dudu #lala .meme .xixi 下面好多级
4. 层级选择器：div > p > a > .lala 只能是下面一级
5. 【注意】select选择器返回永远是列表，须要经过下标提取指定的对象

解析器	使用方法	优点	劣势
Python标准库	`BeautifulSoup(markup, "html.parser")`	Python的内置标准库执行速度适中文档容错能力强	Python 2.7.3 or 3.2.2)前的版本中文档容错能力差
lxml HTML 解析器	`BeautifulSoup(markup, "lxml")`	速度快文档容错能力强	须要安装C语言库
lxml XML 解析器	`BeautifulSoup(markup, ["lxml", "xml"])`css `BeautifulSoup(markup, "xml")`html	速度快惟一支持XML的解析器	须要安装C语言库
html5lib	`BeautifulSoup(markup, "html5lib")`	最好的容错性以浏览器的方式解析文档生成HTML5格式的文档	速度慢不依赖外部扩展

#安装 Beautiful Soup
pip install beautifulsoup4

#安装解析器
Beautiful Soup支持Python标准库中的HTML解析器,还支持一些第三方的解析器,其中一个是 lxml .根据操做系统不一样,能够选择下列方法来安装lxml:

$ apt-get install Python-lxml

$ easy_install lxml

$ pip install lxml

另外一个可供选择的解析器是纯Python实现的 html5lib , html5lib的解析方式与浏览器相同,能够选择下列方法来安装html5lib:

$ apt-get install Python-html5lib

$ easy_install html5lib

$ pip install html5lib

安装lxml

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

#基本使用：容错处理,文档的容错能力指的是在html代码不完整的状况下,使用该模块能够识别该错误。使用BeautifulSoup解析上述代码,可以获得一个 BeautifulSoup 的对象,并能按照标准的缩进格式的结构输出
from bs4 import BeautifulSoup
soup=BeautifulSoup(html_doc,'lxml') #具备容错功能
res=soup.prettify() #处理好缩进，结构化显示
print(res)

基本使用

"""
#遍历文档树：即直接经过标签名字选择，特色是选择速度快，但若是存在多个相同的标签则只返回第一个
#一、用法
#二、获取标签的名称
#三、获取标签的属性
#四、获取标签的内容
#五、嵌套选择
#六、子节点、子孙节点
#七、父节点、祖先节点
#八、兄弟节点
"""
#遍历文档树：即直接经过标签名字选择，特色是选择速度快，但若是存在多个相同的标签则只返回第一个
html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p id="my p" class="title"><b id="bbb" class="boldest">The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

#一、用法
from bs4 import BeautifulSoup
soup=BeautifulSoup(html_doc,'lxml')
# soup=BeautifulSoup(open('a.html'),'lxml')

print(soup.p) #存在多个相同的标签则只返回第一个
print(soup.a) #存在多个相同的标签则只返回第一个

#二、获取标签的名称
print(soup.p.name)

#三、获取标签的属性
print(soup.p.attrs)

#四、获取标签的内容
print(soup.p.string) # p下的文本只有一个时，取到，不然为None
print(soup.p.strings) #拿到一个生成器对象, 取到p下全部的文本内容
print(soup.p.text) #取到p下全部的文本内容
for line in soup.stripped_strings: #去掉空白
    print(line)


'''
若是tag包含了多个子节点,tag就没法肯定 .string 方法应该调用哪一个子节点的内容, .string 的输出结果是 None，若是只有一个子节点那么就输出该子节点的文本，好比下面的这种结构，soup.p.string 返回为None,但soup.p.strings就能够找到全部文本
<p id='list-1'>
    哈哈哈哈
    <a class='sss'>
        <span>
            <h1>aaaa</h1>
        </span>
    </a>
    <b>bbbbb</b>
</p>
'''

#五、嵌套选择
print(soup.head.title.string)
print(soup.body.a.string)


#六、子节点、子孙节点
print(soup.p.contents) #p下全部子节点
print(soup.p.children) #获得一个迭代器,包含p下全部子节点

for i,child in enumerate(soup.p.children):
    print(i,child)

print(soup.p.descendants) #获取子孙节点,p下全部的标签都会选择出来
for i,child in enumerate(soup.p.descendants):
    print(i,child)

#七、父节点、祖先节点
print(soup.a.parent) #获取a标签的父节点
print(soup.a.parents) #找到a标签全部的祖先节点，父亲的父亲，父亲的父亲的父亲...


#八、兄弟节点
print('=====>')
print(soup.a.next_sibling) #下一个兄弟
print(soup.a.previous_sibling) #上一个兄弟

print(list(soup.a.next_siblings)) #下面的兄弟们=>生成器对象
print(soup.a.previous_siblings) #上面的兄弟们=>生成器对象

遍历文档树

3、搜索文档树

一、五种过滤器

#搜索文档树：BeautifulSoup定义了不少搜索方法,这里着重介绍2个: find() 和 find_All() .其它方法的参数和用法相似
html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p id="my p" class="title"><b id="bbb" class="boldest">The Dormouse's story</b>
</p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""


from bs4 import BeautifulSoup
soup=BeautifulSoup(html_doc,'lxml')

#一、五种过滤器: 字符串、正则表达式、列表、True、方法
#1.一、字符串：即标签名
print(soup.find_All('b'))

#1.二、正则表达式
import re
print(soup.find_All(re.compile('^b'))) #找出b开头的标签，结果有body和b标签

#1.三、列表：若是传入列表参数,Beautiful Soup会将与列表中任一元素匹配的内容返回.下面代码找到文档中全部<a>标签和<b>标签:
print(soup.find_All(['a','b']))

#1.四、True：能够匹配任何值,下面代码查找到全部的tag,可是不会返回字符串节点
print(soup.find_All(True))
for tag in soup.find_All(True):
    print(tag.name)

#1.五、方法:若是没有合适过滤器,那么还能够定义一个方法,方法只接受一个元素参数 ,若是这个方法返回 True 表示当前元素匹配而且被找到,若是不是则反回 False
def has_class_but_no_id(tag):
    return tag.has_attr('class') and not tag.has_attr('id')

print(soup.find_All(has_class_but_no_id))

View Code

二、find_All( name , attrs , recursive , text , **kwargs )

#二、find_All( name , attrs , recursive , text , **kwargs )
#2.一、name: 搜索name参数的值可使任一类型的 过滤器 ,字符窜,正则表达式,列表,方法或是 True .
print(soup.find_All(name=re.compile('^t')))

#2.二、keyword: key=value的形式，value能够是过滤器：字符串 , 正则表达式 , 列表, True .
print(soup.find_All(id=re.compile('my')))
print(soup.find_All(href=re.compile('lacie'),id=re.compile('\d'))) #注意类要用class_
print(soup.find_All(id=True)) #查找有id属性的标签

# 有些tag属性在搜索不能使用,好比HTML5中的 data-* 属性:
data_soup = BeautifulSoup('<div data-foo="value">foo!</div>','lxml')
# data_soup.find_All(data-foo="value") #报错：SyntaxError: keyword can't be an expression
# 可是能够经过 find_All() 方法的 attrs 参数定义一个字典参数来搜索包含特殊属性的tag:
print(data_soup.find_All(attrs={"data-foo": "value"}))
# [<div data-foo="value">foo!</div>]

#2.三、按照类名查找，注意关键字是class_，class_=value,value能够是五种选择器之一
print(soup.find_All('a',class_='sister')) #查找类为sister的a标签
print(soup.find_All('a',class_='sister ssss')) #查找类为sister和sss的a标签，顺序错误也匹配不成功
print(soup.find_All(class_=re.compile('^sis'))) #查找类为sister的全部标签

#2.四、attrs
print(soup.find_All('p',attrs={'class':'story'}))

#2.五、text: 值能够是：字符，列表，True，正则
print(soup.find_All(text='Elsie'))
print(soup.find_All('a',text='Elsie'))

#2.六、limit参数:若是文档树很大那么搜索会很慢.若是咱们不须要所有结果,可使用 limit 参数限制返回结果的数量.效果与SQL中的limit关键字相似,当搜索到的结果数量达到 limit 的限制时,就中止搜索返回结果
print(soup.find_All('a',limit=2))

#2.七、recursive:调用tag的 find_All() 方法时,Beautiful Soup会检索当前tag的全部子孙节点,若是只想搜索tag的直接子节点,可使用参数 recursive=False .
print(soup.html.find_All('a'))
print(soup.html.find_All('a',recursive=False))

'''
像调用 find_All() 同样调用tag
find_All() 几乎是Beautiful Soup中最经常使用的搜索方法,因此咱们定义了它的简写方法. BeautifulSoup 对象和 tag 对象能够被看成一个方法来使用,这个方法的执行结果与调用这个对象的 find_All() 方法相同,下面两行代码是等价的:
soup.find_All("a")
soup("a")
这两行代码也是等价的:
soup.title.find_All(text=True)
soup.title(text=True)
'''

View Code

三、find( name , attrs , recursive , text , **kwargs )

#三、find( name , attrs , recursive , text , **kwargs )
find_All() 方法将返回文档中符合条件的全部tag,尽管有时候咱们只想获得一个结果.好比文档中只有一个<body>标签,那么使用 find_All() 方法来查找<body>标签就不太合适, 使用 find_All 方法并设置 limit=1 参数不如直接使用 find() 方法.下面两行代码是等价的:

soup.find_All('title', limit=1)
# [<title>The Dormouse's story</title>]
soup.find('title')
# <title>The Dormouse's story</title>

惟一的区别是 find_All() 方法的返回结果是值包含一个元素的列表,而 find() 方法直接返回结果.
find_All() 方法没有找到目标是返回空列表, find() 方法找不到目标时,返回 None .
print(soup.find("nosuchtag"))
# None

soup.head.title 是 tag的名字 方法的简写.这个简写的原理就是屡次调用当前tag的 find() 方法:

soup.head.title
# <title>The Dormouse's story</title>
soup.find("head").find("title")
# <title>The Dormouse's story</title>

View Code

四、其余方法

#见官网:https://www.crummy.com/software/BeautifulSoup/bs4/doc/index.zh.html#find-parents-find-parent

View Code

五、CSS选择器

#该模块提供了select方法来支持css,详见官网:https://www.crummy.com/software/BeautifulSoup/bs4/doc/index.zh.html#id37
html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title">
    <b>The Dormouse's story</b>
    Once upon a time there were three little sisters; and their names were
    <a href="http://example.com/elsie" class="sister" id="link1">
        <span>Elsie</span>
    </a>
    <a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and
    <a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;
    <div class='panel-1'>
        <ul class='list' id='list-1'>
            <li class='element'>Foo</li>
            <li class='element'>Bar</li>
            <li class='element'>Jay</li>
        </ul>
        <ul class='list list-small' id='list-2'>
            <li class='element'><h1 class='yyyy'>Foo</h1></li>
            <li class='element xxx'>Bar</li>
            <li class='element'>Jay</li>
        </ul>
    </div>
    and they lived at the bottom of a well.
</p>
<p class="story">...</p>
"""
from bs4 import BeautifulSoup
soup=BeautifulSoup(html_doc,'lxml')

#一、CSS选择器
print(soup.p.select('.sister'))
print(soup.select('.sister span'))

print(soup.select('#link1'))
print(soup.select('#link1 span'))

print(soup.select('#list-2 .element.xxx'))

print(soup.select('#list-2')[0].select('.element')) #能够一直select,但其实不必,一条select就能够了

# 二、获取属性
print(soup.select('#list-2 h1')[0].attrs)

# 三、获取内容
print(soup.select('#list-2 h1')[0].get_text())

View Code

4、修改文档树

修改文档树点击

5、总结

推荐使用lxml解析库
讲了三种选择器:标签选择器,find与find_All，css选择器
1. 标签选择器筛选功能弱,可是速度快
2. 建议使用find,find_All查询匹配单个结果或者多个结果
3. 若是对css选择器很是熟悉建议使用select
记住经常使用的获取属性attrs和文本值get_text()的方法

"""
数据解析：
- 1.指定url
- 2.发起请求
- 3.获取页面数据
- 4.数据解析
- 5.进行持久化存储
三种数据解析方式：
- 正则
- bs4
- xpath

使用正则对糗事百科中的图片数据进行解析和下载
<div class="thumb">

<a href="/article/121159481" target="_blank">
<img src="//pic.qiushibaike.com/system/pictures/12115/121159481/medium/CIYK4P1D4DKSBY4L.jpg" alt="不是臭咸鱼吗">
</a>

</div>
"""

import requests
import re
import os
#指定url
url = 'https://www.qiushibaike.com/pic/'
headers={
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_12_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/69.0.3497.100 Safari/537.36',
    }
#发起请求
response = requests.get(url=url,headers=headers)

#获取页面数据
page_text = response.text

#数据解析(该列表中存储的就是当前页面源码中全部的图片连接)
img_list = re.findall('<div class="thumb">.*?<img src="(.*?)".*?>.*?</div>',page_text,re.S)

#建立一个存储图片数据的文件夹
if not os.path.exists('./imgs'):
    os.mkdir('imgs')
for url in img_list:
    #将图片的url进行拼接，拼接成一个完成的url
    img_url = 'https:' + url
    #持久化存储：存储的是图片的数据，并非url。
    #获取图片二进制的数据值
    img_data = requests.get(url=img_url,headers=headers).content
    imgName = url.split('/')[-1]
    imgPath = 'imgs/'+imgName
    with open(imgPath,'wb') as fp:
        fp.write(img_data)
        print(imgName+'写入成功')


"""
使用xpath对段子网中的段子内容和标题进行解析，持久化存储
"""
import requests
from lxml import etree

#1.指定url
url = 'https://ishuo.cn/joke'
#2.发起请求
headers={
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_12_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/69.0.3497.100 Safari/537.36',
    }
response = requests.get(url=url,headers=headers)
#3.获取页面内容
page_text = response.text
#4.数据解析
tree = etree.HTML(page_text)
#获取全部的li标签（段子内容和标题都被包含在li标签中）
li_list = tree.xpath('//div[@id="list"]/ul/li')
#注意：Element类型的对象能够继续调用xpath函数，对该对象表示的局部内容进行指定内容的解析
fp = open('./duanzi.txt','w',encoding='utf-8')
for li in li_list:
    content = li.xpath('./div[@class="content"]/text()')[0]
    title = li.xpath('./div[@class="info"]/a/text()')[0]
    #5.持久化
    fp.write(title+":"+content+"\n\n")
    print('一条数据写入成功')


"""
需求：爬取古诗文网中三国小说里的标题和内容
"""
import requests
from bs4 import BeautifulSoup

url = 'http://www.shicimingju.com/book/sanguoyanyi.html'
headers={
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_12_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/69.0.3497.100 Safari/537.36',
    }


#根据url获取页面内容中指定的标题所对应的文章内容
def get_content(url):
    content_page = requests.get(url=url,headers=headers).text
    soup = BeautifulSoup(content_page,'lxml')
    div = soup.find('div',class_='chapter_content')
    return div.text

page_text = requests.get(url=url,headers=headers).text

#数据解析
soup = BeautifulSoup(page_text,'lxml')
#a_list列表中存储的是一系列的a标签对象
a_list = soup.select('.book-mulu > ul > li > a')
#type(a_list[0])
#注意：Tag类型的对象能够继续调用响应的解析属性和方法进行局部数据的解析

fp = open('./sanguo.txt','w',encoding='utf-8')
for a in a_list:
    #获取了章节的标题
    title = a.string
    content_url = 'http://www.shicimingju.com'+a['href']
    print(content_url)
    #获取章节的内容
    content = get_content(content_url)
    fp.write(title+':'+content+"\n\n\n")
    print('写入一个章节内容')

三种数据解析方式

pyquery模块

'''
强大而又灵活的网页解析库,若是你以为正则写起来太麻烦,若是你以为beutifulsoup
语法太难记,若是你熟悉jquery的语法,那么pyquery是最佳选择


安装pyquery
pip3 install pyquery
'''

html='''
</div><div class="account-signin">
    <ul class="navigation menu" aria-label="Social Media Navigation">
        哈哈哈
        <li class="tier-1 last" aria-haspopup="true">

            <a href="/accounts/login/" title="Sign Up or Sign In to Python.org">Sign In</a>
            <ul class="subnav menu">
                <li class="tier-2 element-1" role="treeitem"><a href="/accounts/signup/">Sign Up / Register</a></li>
                <li class="tier-2 element-2" role="treeitem"><a href="/accounts/login/">Sign In</a></li>
            </ul>

        </li>
    </ul>
</div>
'''


#用法:

#1===========>初始化
#===>字符串初始化
# from pyquery import PyQuery as pq
# doc=pq(html)
# print(doc('.tier-2')) #默认就是css选择器

#===>url初始化
# from pyquery import PyQuery as pq
# doc=pq(url='http://www.baidu.com')
# print(doc('head'))

#===>文件初始化
# from pyquery import PyQuery as pq
# doc=pq(filename='demo.html')
# print(doc('li'))


#2===========>基本css选择器
from pyquery import PyQuery as pq
doc=pq(html)
# print(doc('.tier-2')) #默认就是css选择器

#查找元素

#子元素
# print(doc('li').find('li')) #这里的find是查找全部,可是不必定是直接子元素
# print('==>',doc('li').children('li')) #查找直接子元素


#父元素
# print(doc('.tier-2').parent())

#祖先元素:爹,爹的爹
# print(doc('.tier-2').parents())
# print(doc('.tier-2').parents('.account-signin')) #从祖先里筛选

#先补充:并列选择
# print(doc('.tier-1 .tier-2'))
# print(doc('.tier-1 .tier-2.element-1'))

#兄弟元素
# print(doc('.tier-2.element-1').siblings())
# print(doc('.tier-2.element-1').siblings('li a'))







#3===========>遍历

# lis=doc('li').items()
# print(lis)
#
# for i,j in enumerate(lis):
#     print(i,j)

#4===========>获取属性
# print(doc('li').attr('class'))
# print(doc('a').attr.href)


# 5===========>获取文本
# print(doc('a').text())

#6===========>获取html
# print(doc('.subnav.menu'))
# print(doc('.subnav.menu').html())


#7===========>DOM
#addclass,removeclass
# tag=doc('.subnav.menu')
# print(tag)
#
# tag.addClass('active')
# print(tag)
#
# tag.removeClass('active')
# print(tag)


# tag=doc('.tier-2.element-1 a')
# tag.attr('name','link')
# tag.css('font-size','14px')
# print(tag)


tag=doc('.navigation.menu')
# print(tag.text()) #获取的是tag下全部的文本,

tag.find('li').remove()
print(tag.text()) #若是指向获取url下的那个"哈哈哈",则须要先删除li

#8===========>pyquery官网


# http://pyquery.readthedocs.io/en.latest/api.html


#9===========>伪类选择器

print(doc('li:first-child')) #选择li标签的第一个
print(doc('li:last-child')) #选择li标签的最后一个
print(doc('li:nth-child(2)')) #选择li标签的第2个
print(doc('li:gt(2)')) #选择li标签第2个之后的
print(doc('li:nth-child(2n)')) #选择li标签的偶数标签
print(doc('li:nth-child(2n+1)')) #选择li标签的奇数标签
print(doc('li:contains(second)')) #选择li标签中包含second文本的标签

#更多css选择器能够查看
# http://www.w3school.com.cn/css/index.asp

#官网:http://pyquery.readthedocs.io/

View Code