python使用lxml和xpath解析html表上的特定数据

<html> table class="clCommonGrid" cellspacing="0"> <thead> <tr> <td colspan="3">Kommande matcher</td> </tr> <tr> <th style="width:1%;">Tid</th> <th style="width:69%;">Match</th> <th style="width:30%;">Arena</th> </tr> </thead> <tbody class="clGrid"> <tr class="clTrOdd"> <td nowrap="nowrap" class="no-line-through"> <span class="matchTid"><span>2014-09-26 19:30</span></span> </td> <td><a href="?scr=result&fmid=2669197">Guldhedens IK - IF Warta</a></td> <td><a href="?scr=venue&faid=847">Guldheden Södra 1 Konstgräs</a> </td> </tr> <tr class="clTrEven"> <td nowrap="nowrap" class="no-line-through"> <span class="matchTid"><span>2014-09-26 13:00</span></span> </td> <td><a href="?scr=result&fmid=2669176">Romelanda UF - IK Virgo</a></td> <td><a href="?scr=venue&faid=941">Romevi 1 Gräs</a> </td> </tr> <tr class="clTrOdd"> <td nowrap="nowrap" class="no-line-through"> <span class="matchTid"><span>2014-09-27 13:00</span></span> </td> <td><a href="?scr=result&fmid=2669167">Kode IF - IK Kongahälla</a></td> <td><a href="?scr=venue&faid=912">Kode IP 1 Gräs</a> </td> </tr> <tr class="clTrEven"> <td nowrap="nowrap" class="no-line-through"> <span class="matchTid"><span>2014-09-27 14:00</span></span> </td> <td><a href="?scr=result&fmid=2669147">Floda BoIF - Partille IF FK </a></td> <td><a href="?scr=venue&faid=218">Flodala IP 1</a> </td> </tr> </tbody> </table> </html>

2条回答

网友

1楼 · 编辑于 2024-05-14 09:03:27

如果我能理解你，试试这样的方法：

import lxml.html
url = "http://gbgfotboll.se/information/?scr=table&ftid=51168"
html = lxml.html.parse(url)
for i in range(12):
    xpath1 = ".//*[@id='content-primary']/table[3]/tbody/tr[%d]/td[1]/span/span//text()" %(i+1)
    xpath2 = ".//*[@id='content-primary']/table[3]/tbody/tr[%d]/td[2]/a/text()" %(i+1)
    print html.xpath(xpath1)[1], html.xpath(xpath2)[0]

我知道这是脆弱的，有更好的解决办法，但它是有效的。；）

编辑：
使用BeautifulSoup的更好方法：

^{pr2}$

编辑2: 页面没有响应，但应该可以：

from bs4 import BeautifulSoup
import requests

respond = requests.get("http://gbgfotboll.se/information/?scr=table&ftid=51168")
soup = BeautifulSoup(respond.text)
l = soup.find_all('table')
t = l[2].find_all('tr')
time = ""
for i in t:
    try:
        dateTime = i.find('span').get_text()
        teamName = i.find('a').get_text()
        if time == dateTime[:-5]:
            print dateTime[-5,], teamName
        else:
            print dateTime, teamName
            time = dateTime[:-5]
    except AttributeError:
        pass

lxml公司：

import lxml.html
url = "http://gbgfotboll.se/information/?scr=table&ftid=51168"
html = lxml.html.parse(url)
dateTemp = ""
for i in range(12):
    xpath1 = ".//*[@id='content-primary']/table[3]/tbody/tr[%d]/td[1]/span/span//      text()" %(i+1)
    xpath2 = ".//*[@id='content-primary']/table[3]/tbody/tr[%d]/td[2]/a/text()" %(i+1)
    time = html.xpath(xpath1)[1]
    date = html.xpath(xpath1)[0]
    teamName = html.xpath(xpath2)[0]
    if date == dateTemp:
        print time, teamName
    else:
        print date, time, teamName

网友

2楼 · 编辑于 2024-05-14 09:03:27

所以多亏了@CodeNinja的帮助，我才稍微调整了一下，才得到我想要的东西。我导入时间来获取运行代码的日期。不管怎样，这是我想要的代码。谢谢你的帮助！！在

import lxml.html
import time
url = "http://gbgfotboll.se/information/?scr=table&ftid=51168"
html = lxml.html.parse(url)
currentDate = (time.strftime("%Y-%m-%d"))
for i in range(12):
    xpath1 = ".//*[@id='content-primary']/table[3]/tbody/tr[%d]/td[1]/span/span//text()" %(i+1)
    xpath2 = ".//*[@id='content-primary']/table[3]/tbody/tr[%d]/td[2]/a/text()" %(i+1)
    time = html.xpath(xpath1)[1]
    date = html.xpath(xpath1)[0]
    teamName = html.xpath(xpath2)[0]
    if date == currentDate:
        print time, teamName

所以这里是如何正确操作的最终版本。这将解析它拥有的所有表行，而不在for循环中使用“range”。我从我的另一个帖子得到了这个答案：Iterate through all the rows in a table using python lxml xpath

^{pr2}$

相关问题更多 >

编程相关推荐

热门问题

热门文章