如何在python中将查询提交到.aspx页面

3条回答

网友

1楼 · 编辑于 2024-06-12 13:35:05

Selenium是用于此类任务的一个很好的工具。可以指定要输入的表单值，并在几行python代码中将响应页的html作为字符串检索。使用Selenium，您可能不必手动模拟一个有效的post请求及其所有隐藏变量，正如我在多次尝试后发现的那样。

网友

2楼 · 编辑于 2024-06-12 13:35:05

作为概述，您需要执行四项主要任务：

向网站提交请求
从站点检索响应
分析这些响应
使用与导航相关的参数（到结果列表中的“下一页”）在上面的任务中迭代一些逻辑

http请求和响应处理是使用Python标准库的urllib和urllib2中的方法和类完成的。html页面的解析可以使用Python的标准库HTMLParser或其他模块（如Beautiful Soup）完成

下面的代码片段演示了在问题中指定的站点请求和接收搜索的过程。这个站点是由ASP驱动的，因此我们需要确保发送几个表单字段，其中一些具有“可怕”的值，因为ASP逻辑在某种程度上使用这些值来维护状态和验证请求。真的很顺从。请求必须使用http POST方法发送，因为这是此ASP应用程序所期望的。主要的困难在于识别表单字段和ASP期望的相关值（使用Python获取页面是最简单的部分）。

这段代码是功能性的，或者更准确地说，was是功能性的，直到我删除了大部分VSTATE值，并且可能通过添加注释引入了一两个拼写错误。

import urllib
import urllib2

uri = 'http://legistar.council.nyc.gov/Legislation.aspx'

#the http headers are useful to simulate a particular browser (some sites deny
#access to non-browsers (bots, etc.)
#also needed to pass the content type. 
headers = {
    'HTTP_USER_AGENT': 'Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.0.13) Gecko/2009073022 Firefox/3.0.13',
    'HTTP_ACCEPT': 'text/html,application/xhtml+xml,application/xml; q=0.9,*/*; q=0.8',
    'Content-Type': 'application/x-www-form-urlencoded'
}

# we group the form fields and their values in a list (any
# iterable, actually) of name-value tuples.  This helps
# with clarity and also makes it easy to later encoding of them.

formFields = (
   # the viewstate is actualy 800+ characters in length! I truncated it
   # for this sample code.  It can be lifted from the first page
   # obtained from the site.  It may be ok to hardcode this value, or
   # it may have to be refreshed each time / each day, by essentially
   # running an extra page request and parse, for this specific value.
   (r'__VSTATE', r'7TzretNIlrZiKb7EOB3AQE ... ...2qd6g5xD8CGXm5EftXtNPt+H8B'),

   # following are more of these ASP form fields
   (r'__VIEWSTATE', r''),
   (r'__EVENTVALIDATION', r'/wEWDwL+raDpAgKnpt8nAs3q+pQOAs3q/pQOAs3qgpUOAs3qhpUOAoPE36ANAve684YCAoOs79EIAoOs89EIAoOs99EIAoOs39EIAoOs49EIAoOs09EIAoSs99EI6IQ74SEV9n4XbtWm1rEbB6Ic3/M='),
   (r'ctl00_RadScriptManager1_HiddenField', ''), 
   (r'ctl00_tabTop_ClientState', ''), 
   (r'ctl00_ContentPlaceHolder1_menuMain_ClientState', ''),
   (r'ctl00_ContentPlaceHolder1_gridMain_ClientState', ''),

   #but then we come to fields of interest: the search
   #criteria the collections to search from etc.
                                                       # Check boxes  
   (r'ctl00$ContentPlaceHolder1$chkOptions$0', 'on'),  # file number
   (r'ctl00$ContentPlaceHolder1$chkOptions$1', 'on'),  # Legislative text
   (r'ctl00$ContentPlaceHolder1$chkOptions$2', 'on'),  # attachement
                                                       # etc. (not all listed)
   (r'ctl00$ContentPlaceHolder1$txtSearch', 'york'),   # Search text
   (r'ctl00$ContentPlaceHolder1$lstYears', 'All Years'),  # Years to include
   (r'ctl00$ContentPlaceHolder1$lstTypeBasic', 'All Types'),  #types to include
   (r'ctl00$ContentPlaceHolder1$btnSearch', 'Search Legislation')  # Search button itself
)

# these have to be encoded    
encodedFields = urllib.urlencode(formFields)

req = urllib2.Request(uri, encodedFields, headers)
f= urllib2.urlopen(req)     #that's the actual call to the http site.

# *** here would normally be the in-memory parsing of f 
#     contents, but instead I store this to file
#     this is useful during design, allowing to have a
#     sample of what is to be parsed in a text editor, for analysis.

try:
  fout = open('tmp.htm', 'w')
except:
  print('Could not open output file\n')

fout.writelines(f.readlines())
fout.close()

这是关于获得初始页面的内容。如上所述，然后需要解析页面，即找到感兴趣的部分并酌情收集它们，然后将它们存储到file/database/wherever。这项工作可以通过很多方式来完成：使用html解析器或XSLT类型的技术（实际上是在将html解析为xml之后），甚至对于粗糙的作业，也可以使用简单的正则表达式。此外，通常提取的项目之一是“下一个信息”，即排序的链接，可以在对服务器的新请求中使用，以获取后续页面。

这应该会让您大致了解“长手”html抓取是什么意思。还有很多其他的方法，比如专用的实用程序、Mozilla（FireFox）GreaseMonkey插件中的脚本、XSLT。。。

网友

3楼 · 编辑于 2024-06-12 13:35:05

大多数ASP.NET站点（包括您引用的站点）实际上会使用HTTP post谓词而不是GET谓词将查询发回给自己。这就是为什么URL没有像您所注意到的那样改变。

您需要做的是查看生成的HTML并捕获它们的所有表单值。请确保捕获所有表单值，因为其中一些值用于页面验证，没有它们，您的POST请求将被拒绝。

除了验证，ASPX页面在抓取和发布方面与其他web技术没有区别。

相关问题更多 >

编程相关推荐

热门问题

热门文章