久久久久久久视色,久久电影免费精品,中文亚洲欧美乱码在线观看,在线免费播放AV片

<center id="vfaef"><input id="vfaef"><table id="vfaef"></table></input></center>

<p id="vfaef"><kbd id="vfaef"></kbd></p>

<pre id="vfaef"><u id="vfaef"></u></pre>

<thead id="vfaef"><input id="vfaef"></input></thead>

當前位置：站長資訊網(wǎng) > 編程知識 > 正文

Python之Spider

2020-10-13 分類：編程知識閱讀(1247) 評論(0)

今天python視頻教程欄目為大家介紹Python的Spider （爬蟲）相關(guān)知識。

Python之Spider

一、網(wǎng)絡(luò)爬蟲

網(wǎng)絡(luò)爬蟲又被稱為網(wǎng)絡(luò)蜘蛛，我們可以把互聯(lián)網(wǎng)想象成一個蜘蛛網(wǎng)，每一個網(wǎng)站都是一個節(jié)點，我們可以使用一只蜘蛛去各個網(wǎng)頁抓取我們想要的資源。舉一個最簡單的例子，你在百度和谷歌中輸入‘Python'，會有大量和Python相關(guān)的網(wǎng)頁被檢索出來，百度和谷歌是如何從海量的網(wǎng)頁中檢索出你想要的資源，他們靠的就是派出大量蜘蛛去網(wǎng)頁上爬取，檢索關(guān)鍵字，建立索引數(shù)據(jù)庫，經(jīng)過復雜的排序算法，結(jié)果按照搜索關(guān)鍵字相關(guān)度的高低展現(xiàn)給你。

千里之行，始于足下，我們從最基礎(chǔ)的開始學習如何寫一個網(wǎng)絡(luò)爬蟲，實現(xiàn)語言使用Python。

二、Python如何訪問互聯(lián)網(wǎng)

想要寫網(wǎng)絡(luò)爬蟲，第一步是訪問互聯(lián)網(wǎng)，Python如何訪問互聯(lián)網(wǎng)呢？

在Python中，我們使用urllib包訪問互聯(lián)網(wǎng)。（在Python3中，對這個模塊做了比較大的調(diào)整，以前有urllib和urllib2,在3中對這兩個模塊做了統(tǒng)一合并，稱為urllib包。包下面包含了四個模塊，urllib.request，urllib.error，urllib.parse，urllib.robotparser），目前主要使用的是urllib.request。

我們首先舉一個最簡單的例子，如何獲取獲取網(wǎng)頁的源碼：

import urllib.request response = urllib.request.urlopen('https://docs.python.org/3/') html = response.read()print(html.decode('utf-8'))

三、Python網(wǎng)絡(luò)簡單使用

首先我們用兩個小demo練一下手，一個是使用python代碼下載一張圖片到本地，另一個是調(diào)用有道翻譯寫一個翻譯小軟件。

3.1根據(jù)圖片鏈接下載圖片，代碼如下：

import urllib.request  response = urllib.request.urlopen('http://www.3lian.com/e/ViewImg/index.html?url=http://img16.3lian.com/gif2016/w1/3/d/61.jpg') image = response.read()  with open('123.jpg','wb') as f:     f.write(image)

其中response是一個對象

輸入：response.geturl()

->'http://www.3lian.com/e/ViewImg/index.html?url=http://img16.3lian.com/gif2016/w1/3/d/61.jpg'
輸入：response.info()

-><http.client.HTTPMessage object at 0x10591c0b8>

輸入：print(response.info())

->Content-Type: text/html
Last-Modified: Mon, 27 Sep 2004 01:23:20 GMT
Accept-Ranges: bytes
ETag: "0f4b59230a4c41:0"
Server: Microsoft-IIS/8.0
Date: Sun, 14 Aug 2016 07:16:01 GMT
Connection: close

Content-Length: 2827

輸入：response.getcode()

->200

3.1使用有道詞典實現(xiàn)翻譯功能

我們想實現(xiàn)翻譯功能，我們需要拿到請求鏈接。首先我們需要進入有道首頁，點擊翻譯，在翻譯界面輸入要翻譯的內(nèi)容，點擊翻譯按鈕，就會向服務(wù)器發(fā)起一個請求，我們需要做的就是拿到請求地址和請求參數(shù)。

我在此使用谷歌瀏覽器實現(xiàn)拿到請求地址和請求參數(shù)。首先點擊右鍵，點擊檢查（不同瀏覽器點擊的選項可能不同，同一瀏覽器的不同版本也可能不同），進入圖一所示，從中我們可以拿到請求請求地址和請求參數(shù)，在Header中的Form Data中我們可以拿到請求參數(shù)。

Python之Spider

（圖一）

代碼段如下：

import urllib.requestimport urllib.parse  url = 'http://fanyi.youdao.com/translate?smartresult=dict&smartresult=rule&smartresult=ugc&sessionFrom=dict2.index'data = {} data['type'] = 'AUTO'data['i'] = 'i love you'data['doctype'] = 'json'data['xmlVersion'] = '1.8'data['keyfrom'] = 'fanyi.web'data['ue'] = 'UTF-8'data['action'] = 'FY_BY_CLICKBUTTON'data['typoResult'] = 'true'data = urllib.parse.urlencode(data).encode('utf-8') response = urllib.request.urlopen(url,data) html = response.read().decode('utf-8')print(html)

上述代碼執(zhí)行如下：

{"type":"EN2ZH_CN","errorCode":0,"elapsedTime":0,"translateResult":[[{"src":"i love you","tgt":"我愛你"}]],"smartResult":{"type":1,"entries":["","我愛你。"]}}

對于上述結(jié)果，我們可以看到是一個json串，我們可以對此解析一下，并且對代碼進行完善一下：

import urllib.requestimport urllib.parseimport json  url = 'http://fanyi.youdao.com/translate?smartresult=dict&smartresult=rule&smartresult=ugc&sessionFrom=dict2.index'data = {} data['type'] = 'AUTO'data['i'] = 'i love you'data['doctype'] = 'json'data['xmlVersion'] = '1.8'data['keyfrom'] = 'fanyi.web'data['ue'] = 'UTF-8'data['action'] = 'FY_BY_CLICKBUTTON'data['typoResult'] = 'true'data = urllib.parse.urlencode(data).encode('utf-8') response = urllib.request.urlopen(url,data) html = response.read().decode('utf-8') target = json.loads(html)print(target['translateResult'][0][0]['tgt'])

四、規(guī)避風險

服務(wù)器檢測出請求不是來自瀏覽器，可能會屏蔽掉請求，服務(wù)器判斷的依據(jù)是使用‘User-Agent',我們可以修改改字段的值，來隱藏自己。代碼如下：

import urllib.requestimport urllib.parseimport json  url = 'http://fanyi.youdao.com/translate?smartresult=dict&smartresult=rule&smartresult=ugc&sessionFrom=dict2.index'data = {} data['type'] = 'AUTO'data['i'] = 'i love you'data['doctype'] = 'json'data['xmlVersion'] = '1.8'data['keyfrom'] = 'fanyi.web'data['ue'] = 'UTF-8'data['action'] = 'FY_BY_CLICKBUTTON'data['typoResult'] = 'true'data = urllib.parse.urlencode(data).encode('utf-8') req = urllib.request.Request(url, data) req.add_header('User-Agent','Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_2) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/52.0.2743.116 Safari/537.36') response = urllib.request.urlopen(url, data) html = response.read().decode('utf-8') target = json.loads(html)print(target['translateResult'][0][0]['tgt'])

View Code

上述做法雖然可以隱藏自己，但是還有很大問題，例如一個網(wǎng)絡(luò)爬蟲下載圖片軟件，在短時間內(nèi)大量下載圖片，服務(wù)器可以可以根據(jù)IP訪問次數(shù)判斷是否是正常訪問。所有上述做法還有很大的問題。我們可以通過兩種做法解決辦法，一是使用延遲，例如5秒內(nèi)訪問一次。另一種辦法是使用代理。

延遲訪問（休眠5秒，缺點是訪問效率低下）：

import urllib.requestimport urllib.parseimport jsonimport timewhile True:     content = input('please input content(input q exit program):')    if content == 'q':        break;      url = 'http://fanyi.youdao.com/translate?smartresult=dict&smartresult=rule&smartresult=ugc&sessionFrom=dict2.index'     data = {}     data['type'] = 'AUTO'     data['i'] = content     data['doctype'] = 'json'     data['xmlVersion'] = '1.8'     data['keyfrom'] = 'fanyi.web'     data['ue'] = 'UTF-8'     data['action'] = 'FY_BY_CLICKBUTTON'     data['typoResult'] = 'true'     data = urllib.parse.urlencode(data).encode('utf-8')     req = urllib.request.Request(url, data)     req.add_header('User-Agent','Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_2) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/52.0.2743.116 Safari/537.36')     response = urllib.request.urlopen(url, data)     html = response.read().decode('utf-8')     target = json.loads(html)    print(target['translateResult'][0][0]['tgt'])     time.sleep(5)

View Code

代理訪問：讓代理訪問資源，然后講訪問到的資源返回。服務(wù)器看到的是代理的IP地址，不是自己地址，服務(wù)器就沒有辦法對你做限制。

步驟：

1，參數(shù)是一個字典｛'類型' : '代理IP：端口號' ｝ //類型是http,https等

proxy_support = urllib.request.ProxyHandler({})

2，定制、創(chuàng)建一個opener

opener = urllib.request.build_opener(proxy_support)

3，安裝opener(永久安裝，一勞永逸)

urllib.request.install_opener(opener)

3，調(diào)用opener（調(diào)用的時候使用）

opener.open(url)

五、批量下載網(wǎng)絡(luò)圖片

圖片下載來源為煎蛋網(wǎng)（http://jandan.net)

圖片下載的關(guān)鍵是找到圖片的規(guī)律，如找到當前頁，每一頁的圖片鏈接，然后使用循環(huán)下載圖片。下面是程序代碼（待優(yōu)化,正則表達式匹配，IP代理）：

import urllib.requestimport osdef url_open(url):     req = urllib.request.Request(url)     req.add_header('User-Agent','Mozilla/5.0')     response = urllib.request.urlopen(req)     html = response.read()    return htmldef get_page(url):     html = url_open(url).decode('utf-8')     a = html.find('current-comment-page') + 23     b = html.find(']',a)    return html[a:b]def find_image(url):     html = url_open(url).decode('utf-8')     image_addrs = []     a = html.find('img src=')    while a != -1:         b = html.find('.jpg',a,a + 150)        if b != -1:             image_addrs.append(html[a+9:b+4])        else:             b = a + 9         a = html.find('img src=',b)    for each in image_addrs:        print(each)    return image_addrsdef save_image(folder,image_addrs):    for each in image_addrs:         filename = each.split('/')[-1]         with open(filename,'wb') as f:             img = url_open(each)             f.write(img)def download_girls(folder = 'girlimage',pages = 20):     os.mkdir(folder)     os.chdir(folder)     url = 'http://jandan.net/ooxx/'     page_num = int(get_page(url))    for i in range(pages):         page_num -= i         page_url = url + 'page-' + str(page_num) + '#comments'         image_addrs = find_image(page_url)         save_image(folder,image_addrs)if __name__ == '__main__':     download_girls()

代碼運行效果如下： Python之Spider

贊(0)

分享到：更多 (0)

標簽：AI app NEC php python 互聯(lián)網(wǎng)+關(guān)鍵字數(shù)據(jù)庫服務(wù)器正則表達式瀏覽器百度谷歌

上一篇
Vue中值得關(guān)注的21個開源項目（推薦）下一篇
Vue.js 學習記錄之一：學習規(guī)劃和了解 Vue.js

相關(guān)推薦
華納云香港高防服務(wù)器150G防御4.6折促銷，低至6888元/月，CN2大帶寬直連清洗，終身循環(huán)折扣
RakSmart服務(wù)器成本優(yōu)化策略
2025年國內(nèi)免費AI工具推薦：文章生成與圖像創(chuàng)作全攻略
自媒體推廣實時監(jiān)控從服務(wù)器帶寬到用戶行為解決方法
站長必讀：從“流量思維”到“IP思維”的品牌升級之路
從流量變現(xiàn)到信任變現(xiàn)：個人站長的私域運營方法論
傳統(tǒng)網(wǎng)站如何借力短視頻？從SEO到“內(nèi)容種草”的轉(zhuǎn)型策略
AI時代，個人站長如何用AI工具實現(xiàn)“一人公司”

熱門標簽
word (73806)互聯(lián)網(wǎng)+ (37958)AI (33722)電腦 (26436)app (22712)php (21416)java (19487)美國 (18366)javaScript (17290)蘋果 (15835)5G (15554)處理器 (12917)華為 (11937)人工智能 (11722)大數(shù)據(jù) (11297)服務(wù)器 (11201)list (10960)營銷 (10842)內(nèi)存 (10003)三星 (9977)微軟 (9923)谷歌 (9687)智能手機 (9681)微信 (9334)騰訊 (9214)電商 (8762)百度 (8573)汽車 (8313)直播 (8252)set (8233)

近期文章

華納云香港高防服務(wù)器150G防御4.6折促銷，低至6888元/月，CN2大帶寬直連清洗，終身循環(huán)折扣

RakSmart服務(wù)器成本優(yōu)化策略

2025年國內(nèi)免費AI工具推薦：文章生成與圖像創(chuàng)作全攻略

自媒體推廣實時監(jiān)控從服務(wù)器帶寬到用戶行為解決方法

站長必讀：從“流量思維”到“IP思維”的品牌升級之路

站長必讀：從“流量思維”到“IP思維”的品牌升級之路

從流量變現(xiàn)到信任變現(xiàn)：個人站長的私域運營方法論

傳統(tǒng)網(wǎng)站如何借力短視頻？從SEO到“內(nèi)容種草”的轉(zhuǎn)型策略

AI時代，個人站長如何用AI工具實現(xiàn)“一人公司”

個人站長消亡論？從“消失”到“重生”的三大破局路徑

2025年5月

一二三四五六日

« 4月

1 2 3 4

5 6 7 8 9 10 11

12 13 14 15 16 17 18

19 20 21 22 23 24 25

26 27 28 29 30 31

網(wǎng)站地圖滬ICP備18035694號-2 滬公網(wǎng)安備31011702889846號

久久久久久久视色,久久电影免费精品,中文亚洲欧美乱码在线观看,在线免费播放AV片

最新国产自产视频在线观看国产黃片曰版很爽免费视频亚洲伊人a和欧美伊人和a 亚洲一区国产美女在线速度快