Python爬蟲庫requests獲取響應內容、響應狀態碼、響應頭

阿新 • • 發佈：2020-01-31

首先在程式中引入Requests模組

import requests

一、獲取不同型別的響應內容

在傳送請求後，伺服器會返回一個響應內容，而且requests通常會自動解碼響應內容

1.文字響應內容

獲取文字型別的響應內容

r = requests.get('https://www.baidu.com')
r.text # 通過文字的形式獲取響應內容

'<!DOCTYPE html>\r\n<!--STATUS OK--><html> <head><meta http-equiv=content-type content=text/html;charset=utf-8><meta http-equiv=X-UA-Compatible content=IE=Edge><meta content=always name=referrer><link rel=stylesheet type=text/css href=https://ss1.bdstatic.com/5eN1bjq8AAUYm2zgoY3K/r/www/cache/bdorz/baidu.min.css><title>ç\x99¾åo|ä¸\x80ä¸\x8bï¼\x8cä½\xa0å°±ç\x9f￥é\x81\x93</title></head> <body link=#0000cc> <div id=wrapper> <div id=head> <div class=head_wrapper> <div class=s_form> <div class=s_form_wrapper> <div id=lg> <img hidefocus=true src=//www.baidu.com/img/bd_logo1.png width=270 height=129> </div> <form id=form name=f action=//www.baidu.com/s class=fm> <input type=hidden name=bdorz_come value=1> <input type=hidden name=ie value=utf-8> <input type=hidden name=f value=8> <input type=hidden name=rsv_bp value=1> <input type=hidden name=rsv_idx value=1> <input type=hidden name=tn value=baidu><span class="bg s_ipt_wr"><input id=kw name=wd class=s_ipt value maxlength=255 autocomplete=off autofocus=autofocus></span><span class="bg s_btn_wr"><input type=submit id=su value=ç\x99¾åo|ä¸\x80ä¸\x8b class="bg s_btn" autofocus></span> </form> </div> </div> <div id=u1> <a href=http://news.baidu.com name=tj_trnews class=mnav>æ\x96°é\x97»</a> <a href=https://www.hao123.com name=tj_trhao123 class=mnav>hao123</a> <a href=http://map.baidu.com name=tj_trmap class=mnav>å\x9c°å\x9b¾</a> <a href=http://v.baidu.com name=tj_trvideo class=mnav>è§\x86é￠\x91</a> <a href=http://tieba.baidu.com name=tj_trtieba class=mnav>è′′å\x90§</a> <noscript> <a href=http://www.baidu.com/bdorz/login.gif?login&tpl=mn&u=http%3A%2F%2Fwww.baidu.com%2f%3fbdorz_come%3d1 name=tj_login class=lb>ç\x99»å½\x95</a> </noscript> <script>document.write(\'<a href="http://www.baidu.com/bdorz/login.gif?login&tpl=mn&u=\'+ encodeURIComponent(window.location.href+ (window.location.search === " rel="external nofollow" " ? "?" : "&")+ "bdorz_come=1")+ \'" name="tj_login" class="lb">ç\x99»å½\x95</a>\');\r\n        </script> <a href=//www.baidu.com/more/ name=tj_briicon class=bri style="display: block;">æ\x9b′å¤\x9aäo§å\x93\x81</a> </div> </div> </div> <div id=ftCon> <div id=ftConw> <p id=lh> <a href=http://home.baidu.com>å\x853äo\x8eç\x99¾åo|</a> <a href=http://ir.baidu.com>About Baidu</a> </p> <p id=cp>©2017Baidu<a href=http://www.baidu.com/duty/>ä½¿ç\x94¨ç\x99¾åo|å\x89\x8då¿\x85èˉ»</a> <a href=http://jianyi.baidu.com/ class=cp-feedback>æ\x84\x8fè§\x81å\x8f\x8dé|\x88</a>äo¬ICPèˉ\x81030173å\x8f· <img src=//www.baidu.com/img/gs.gif> </p> </div> </div> </div> </body> </html>\r\n'

通過encoding來獲取響應內容的編碼以及修改編碼

r.encoding

'ISO-8859-1'

2.二進位制響應內容

r.content # 通過content獲取的內容便是二進位制型別的

3.JSON響應內容

r.json()

4.原始響應內容

r = requests.get('https://www.baidu.com',stream=True)
print(r.raw) # 就是urllib中的HTTPResponse物件
print(r.raw.read(10))

<requests.packages.urllib3.response.HTTPResponse object at 0x00000077940AEEF0>
b'\x1f\x8b\x08\x00\x00\x00\x00\x00\x00\x03'

二、響應狀態碼

獲取響應狀態碼

r = requests.get('https://www.baidu.com')
r.status_code

判斷響應狀態碼

r.status_code == requests.codes.ok

True

當傳送一個錯誤請求時，丟擲異常

bad_r = requests.get('http://httpbin.org/status/404')
print(bad_r.status_code)
bad_r.raise_for_status()

404



---------------------------------------------------------------------------

HTTPError                 Traceback (most recent call last)

<ipython-input-15-9b812f4c5860> in <module>()
   1 bad_r = requests.get('http://httpbin.org/status/404')
   2 print(bad_r.status_code)
----> 3 bad_r.raise_for_status()


D:\Anaconda3\lib\site-packages\requests\models.py in raise_for_status(self)
  926 
  927     if http_error_msg:
--> 928       raise HTTPError(http_error_msg,response=self)
  929 
  930   def close(self):


HTTPError: 404 Client Error: NOT FOUND for url: http://httpbin.org/status/404

三、響應頭

獲取響應頭

r = requests.get('https://www.baidu.com')
r.headers

{'Cache-Control': 'private,no-cache,no-store,proxy-revalidate,no-transform','Connection': 'Keep-Alive','Content-Encoding': 'gzip','Content-Type': 'text/html','Date': 'Mon,23 Jul 2018 09:04:12 GMT','Last-Modified': 'Mon,23 Jan 2017 13:23:51 GMT','Pragma': 'no-cache','Server': 'bfe/1.0.8.18','Set-Cookie': 'BDORZ=27315; max-age=86400; domain=.baidu.com; path=/','Transfer-Encoding': 'chunked'}

獲取響應頭的具體欄位

print(r.headers['Server'])
print(r.headers.get('Server'))

bfe/1.0.8.18
bfe/1.0.8.18

更多關於Python爬蟲庫requestsr的使用方法請檢視下面的相關連結

Python爬蟲庫requests獲取響應內容、響應狀態碼、響應頭

首先在程式中引入Requests模組 import requests 一、獲取不同型別的響應內容在傳送請求後，伺服器會返回一個響應內容，而且requests通常會自動解碼響應內容

使用Python爬蟲庫requests傳送請求、傳遞URL引數、定製headers

首先我們先引入requests模組 import requests 一、傳送請求 r = requests.get(\'https://api.github.com/events\') # GET請求

Python爬蟲庫BeautifulSoup獲取物件(標籤)名,屬性,內容,註釋

一、Tag(標籤)物件 1.Tag物件與XML或HTML原生文件中的tag相同。 from bs4 import BeautifulSoup

python爬蟲開發之使用python爬蟲庫requests，urllib與今日頭條搜尋功能爬取搜尋內容例項

使用python爬蟲庫requests，urllib爬取今日頭條街拍美圖程式碼均有註釋 import re,json,requests,os

使用Python爬蟲庫requests傳送表單資料和JSON資料

匯入Python爬蟲庫Requests import requests 一、傳送表單資料要傳送表單資料，只需要將一個字典傳遞給引數data

python爬蟲開發之使用Python爬蟲庫requests多執行緒抓取貓眼電影TOP100例項

使用Python爬蟲庫requests多執行緒抓取貓眼電影TOP100思路：檢視網頁原始碼抓取單頁內容

常用python爬蟲庫介紹與簡要說明

這個列表包含與網頁抓取和資料處理的Python庫 python網路庫通用 urllib -網路庫(stdlib)。

Python爬蟲庫BeautifulSoup的介紹與簡單使用例項

一、介紹 BeautifulSoup庫是靈活又方便的網頁解析庫，處理高效，支援多種解析器。利用它不用編寫正則表示式即可方便地實現網頁資訊的提取。

使用Python爬蟲庫BeautifulSoup遍歷文件樹並對標籤進行操作詳解

下面就是使用Python爬蟲庫BeautifulSoup對文件樹進行遍歷並對標籤進行操作的例項，都是最基礎的內容

python爬蟲庫scrapy簡單使用例項詳解

最近因為專案需求，需要寫個爬蟲爬取一些題庫。在這之前爬蟲我都是用node或者php寫的。一直聽說python寫爬蟲有一手，便入手了python的爬蟲框架scrapy.

Python爬蟲工具requests-html使用解析

使用Python開發的同學一定聽說過Requsts庫，它是一個用於傳送HTTP請求的測試。如比我們用Python做基於HTTP協議的介面測試，那麼一定會首選Requsts，因為它即簡單又強大。現在作者Kenneth Reitz 又開發了requests-htm

python爬蟲例項之獲取動漫截圖

引言之前有些無聊（呆在家裡實在玩的膩了），然後就去B站看了一些python爬蟲視訊，沒有進行基礎的理論學習，也就是直接開始實戰，感覺跟背公式一樣的進行爬蟲，也算行吧，至少還能爬一些東西，hhh。我今天來分享一

python爬蟲使用requests傳送post請求示例詳解

簡介 HTTP協議規定post提交的資料必須放在訊息主體中，但是協議並沒有規定必須使用什麼編碼方式。服務端通過是根據請求頭中的Content-Type欄位來獲知請求中的訊息主體是用何種方式進行編碼，再對訊息主體進行解析。具

Python爬蟲的requests模組你真的學會了嗎？來看看這些高階用法！

1. 檔案上傳我們知道requests可以模擬提交一些資料。假如有的網站需要上傳檔案，我們也可以用它來實現，這非常簡單，示例如下：

python爬蟲用scrapy獲取影片的例項分析

我們平時生活的娛樂中，看電影是大部分小夥伴都喜歡的事情。周圍的人總會有意無意的在談論，有什麼影片上映，好不好看之類的話題，沒事的時候談論電影是非常不錯的話題。那麼，一些好看的影片如果不去電影院的話，在

python爬蟲關於requests.exceptions.ConnectionError 等問題

技術標籤：python 在爬蟲中報如下的錯誤： requests.exceptions.ConnectionError: (‘Connection aborted.’, RemoteDisconnected(‘Remote end closed connection without response’,))

Python爬蟲教你獲取4K超清桌布圖片，手把手教你跟我一起爬！

本文的文字及圖片來源於網路,僅供學習、交流使用,不具有任何商業用途,版權歸原作者所有,如有問題請及時聯絡我們以作處理

python爬蟲(使用requests)報錯，UnicodeEncodeError: ‘latin-1‘ codec can‘t encode characters in position

技術標籤：爬蟲分割槽python爬蟲post 1、初學爬蟲，在寫爬取拉勾網職位資訊程式時，遇到報錯如下：

Python 爬蟲之設定ip代理，設定User-Agent，設定請求頭，設定post載荷

1、get方式：如何為爬蟲新增ip代理，設定Request header（請求頭） import urllib import urllib.request

python-處理日誌檔案，找出各個介面狀態碼為 200時的平均響應時間

今天又一面試題目，可惜我依舊新手，不熟練，速度太慢背景：需要寫一個方法，處理一個程式的日誌檔案。引數檔名稱日誌檔案的特點是：每一行都是收到的程式請求的記錄每一行的格式是：時間日誌級別介面名稱介面處

Python爬蟲庫requests獲取響應內容、響應狀態碼、響應頭

相關推薦