社区所有版块导航
Python
python开源   Django   Python   DjangoApp   pycharm  
DATA
docker   Elasticsearch  
aigc
aigc   chatgpt  
WEB开发
linux   MongoDB   Redis   DATABASE   NGINX   其他Web框架   web工具   zookeeper   tornado   NoSql   Bootstrap   js   peewee   Git   bottle   IE   MQ   Jquery  
机器学习
机器学习算法  
Python88.com
反馈   公告   社区推广  
产品
短视频  
印度
印度  
Py学习  »  Python

需要了解如何使用python刮取实时流数据的帮助

Joel.M • 7 年前 • 2329 次点击  

我正试着收集这个本地web服务器的当前使用情况。这个数字每秒更新一次由随机数生成器生成的值。

当前时间:07:25:16 UTC

当前日期:2018-11-28 UTC

电流使用:13kw

到目前为止,这就是我对美丽集团的尝试:

import requests
from bs4 import BeatifulSoup
import time

def get_count():
  url = "http://10.0.0.206/apps/cy8ckit_062_demo/main.html"

  # request with fake header, otherwise you will get an 403 HTTP error
  r = request.get(url, headers={'User-Agent': Mozilla/5.0})

while True:
  print(get_count())
  time.sleep(8)

但是,当我运行这个脚本时,我每8秒得到一个输出'none'

以下是Web服务器检查的输出:

当前时间:07:39:42 UTC

当前日期2018-11-28 UTC

电流使用:8kw

我一直在试着遵循这一点: How to scrape real time streaming data with Python?

这是我在尝试@chitown88代码后得到的输出:

Traceback (most recent call last):
  File "C:/seniord/csusite/readweb.py", line 14, in <module>
    soup = BeautifulSoup(r.text, 'html.parser')
NameError: name 'r' is not defined

在尝试了@chitown88中修改后的代码后,我将其作为输出(不显示动态值,但我认为beautifulsoup解决了这个问题):

<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN"
"http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd">

<html lang="en" xml:lang="en" xmlns="http://www.w3.org/1999/xhtml">
<head>
<link href="../../styles/buttons.css" rel="stylesheet" type="text/css"/>
<title>CE222494 PSoC 6 WICED WiFi Demo</title>
<meta content="text/html; charset=utf-8" http-equiv="Content-Type"/>
<script src="../../scripts/general_ajax_script.js" type="text/javascript"></script>
<script type="text/javascript">
    /* <![CDATA[ */
       function reloadData()
       {
         do_ajax('/temp_report.html', ajax_handler);
         timeoutID = setTimeout('reloadData()', 500);
       }
      function ajax_handler( result, data )
      {
        switch( result )
        {
            case AJAX_PARTIAL_PROGRESS:
                break;
            case AJAX_STARTING:
                break;
            case AJAX_FINISHED:
                document.getElementById("currentData").innerHTML = data;
                break;
            case AJAX_NO_BROWSER_SUPPORT:
                document.getElementById("currentData").innerHTML = "Failed - your browser does not support this script";
                break;
            case AJAX_FAILED:
                document.getElementById("currentData").innerHTML = "There was a problem retrieving data";
                break;
        }
      }

    /* ]]> */
    </script>
</head>
<body onload="reloadData()">
<div id="currentData">Retrieving current usage data...
    </div>
</body>
</html>
Python社区是高质量的Python/Django开发社区
本文地址:http://www.python88.com/topic/40207
文章 [ 2 ]  |  最新文章 7 年前
chitown88
Reply   •   1 楼
chitown88    7 年前

你的代码不完整。具体来说,1)您实际上没有使用beautifulsoup来做任何事情,2)您的函数没有返回任何东西,这就是为什么它会打印“none”

import pandas as pd
import bs4 
from requests_html import HTMLSession 
import time

def get_count():

    url = 'http://10.0.0.206/apps/cy8ckit_062_demo/main.html'

    session = HTMLSession()
    r = session.get(url)
    r.html.render(sleep=5,timeout=8)

    soup = bs4.BeautifulSoup(r.text,'html.parser')

    data = soup.findAll('div', {'id':'currentData'})[0]
    temp_data = data.findAll('p')
    current_time = temp_data[0].text
    current_date = temp_data[1].text
    current_usage = temp_data[2].text

    print ('%s\n%s\n%s' %(current_time, current_date, current_usage))



while True:
    get_count()
    time.sleep(8)
ewwink
Reply   •   2 楼
ewwink    7 年前

main.html 是错误的url,它用于显示来自 temp_report.html (阿贾克斯)

import requests
from bs4 import BeatifulSoup
import time

def get_count():
    url = "http://10.0.0.206/temp_report.html
    # or
    # url = "http://10.0.0.206/apps/cy8ckit_062_demo/temp_report.html

    # request with fake header, otherwise you will get an 403 HTTP error
    r = request.get(url, headers={'User-Agent': Mozilla/5.0})
    page_source = r.text
    # print(page_source)

    soup = BeautifulSoup(page_source, 'html.parser')
    print(soup)

    # html_body = soup.find('body') # <body>this_text</body>
    # print(html_body.text) # this_text

    # paragraphs = soup.find_all('p') # <body> <p>p1</p> <p>p2</p> </body>
    # for p in paragraphs:
    #    print(p.text) # p1, p2


while True:
    print(get_count())
    time.sleep(8)