如何用tweepy把所有的tweets放在标签上？

# if this is the first time, creates a new array which # will store max id of the tweets for each keyword if not os.path.isfile("max_ids.npy"): max_ids = np.empty(len(keywords)) # every value is initialized as -1 in order to start from the beginning the first time program run max_ids.fill(-1) else: max_ids = np.load("max_ids.npy") # loads the previous max ids # if there is any new keywords added, extends the max_ids array in order to correspond every keyword if len(keywords) > len(max_ids): new_indexes = np.empty(len(keywords) - len(max_ids)) new_indexes.fill(-1) max_ids = np.append(arr=max_ids, values=new_indexes) count = 0 for i in range(len(keywords)): since_date="2015-01-01" sinceId = None tweetCount = 0 maxTweets = 5000000000000000000000 # maximum tweets to find per keyword tweetsPerQry = 100 searchQuery = "#{0}".format(keywords[i]) while tweetCount < maxTweets: if max_ids[i] < 0: if (not sinceId): new_tweets = api.search(q=searchQuery, count=tweetsPerQry) else: new_tweets = api.search(q=searchQuery, count=tweetsPerQry, since_id=sinceId) else: if (not sinceId): new_tweets = api.search(q=searchQuery, count=tweetsPerQry, max_id=str(max_ids - 1)) else: new_tweets = api.search(q=searchQuery, count=tweetsPerQry, max_id=str(max_ids - 1), since_id=sinceId) if not new_tweets: print("Keyword: {0} No more tweets found".format(searchQuery)) break for tweet in new_tweets: count += 1 print(count) file_write.write( . . . ) item = { . . . . . } # instead of using mongo's id for _id, using tweet's id raw_data = tweet._json raw_data["_id"] = tweet.id raw_data.pop("id", None) try: db["Tweets"].insert_one(item) except pymongo.errors.DuplicateKeyError as e: print("Already exists in 'Tweets' collection.") try: db["RawTweets"].insert_one(raw_data) except pymongo.errors.DuplicateKeyError as e: print("Already exists in 'RawTweets' collection.") tweetCount += len(new_tweets) print("Downloaded {0} tweets".format(tweetCount)) max_ids[i] = new_tweets[-1].id np.save(arr=max_ids, file="max_ids.npy") # saving in order to continue mining from where left next time program run

3条回答

网友

1楼 · 编辑于 2024-05-23 18:21:13

对不起，我不能回答，太长时间了。：）

当然：）检查以下示例：高级搜索#数据关键字2015年5月-2016年7月得到这个url:https://twitter.com/search?l=&q=%23data%20since%3A2015-05-01%20until%3A2016-07-31&src=typd

session = requests.session()
keyword = 'data'
date1 = '2015-05-01'
date2 = 2016-07-31
session.get('https://twitter.com/search?l=&q=%23+keyword+%20since%3A+date1+%20until%3A+date2&src=typd', streaming = True)

现在我们有了所有请求的tweets，可能你对“分页”有问题分页url->

https://twitter.com/i/search/timeline?vertical=news&q=%23data%20since%3A2015-05-01%20until%3A2016-07-31&src=typd&include_available_features=1&include_entities=1&max_position=TWEET-759522481271078912-759538448860581892-BD1UO2FFu9QAAAAAAAAETAAAAAcAAAASAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA&reset_error_state=false

也许你可以随机输入一个tweet id，或者你可以首先解析，或者从twitter请求一些数据。这是可以做到的。

使用Chrome的“网络”选项卡查找所有请求的信息：）

网友

2楼 · 编辑于 2024-05-23 18:21:13

看看这个：https://tweepy.readthedocs.io/en/v3.5.0/cursor_tutorial.html

试试这个：

import tweepy

auth = tweepy.OAuthHandler(CONSUMER_TOKEN, CONSUMER_SECRET)
api = tweepy.API(auth)

for tweet in tweepy.Cursor(api.search, q='#python', rpp=100).items():
    # Do something
    pass

在你的案例中，你有最多的tweets要获取，因此根据链接的教程，你可以：

import tweepy

MAX_TWEETS = 5000000000000000000000

auth = tweepy.OAuthHandler(CONSUMER_TOKEN, CONSUMER_SECRET)
api = tweepy.API(auth)

for tweet in tweepy.Cursor(api.search, q='#python', rpp=100).items(MAX_TWEETS):
    # Do something
    pass

如果你想在给定的ID之后发送tweets，你也可以传递这个参数。

网友

3楼 · 编辑于 2024-05-23 18:21:13

这个密码对我有效。

import tweepy
import pandas as pd
import os

#Twitter Access
auth = tweepy.OAuthHandler( 'xxx','xxx')
auth.set_access_token('xxx-xxx','xxx')
api = tweepy.API(auth,wait_on_rate_limit = True)

df = pd.DataFrame(columns=['text', 'source', 'url'])
msgs = []
msg =[]

for tweet in tweepy.Cursor(api.search, q='#bmw', rpp=100).items(10):
    msg = [tweet.text, tweet.source, tweet.source_url] 
    msg = tuple(msg)                    
    msgs.append(msg)

df = pd.DataFrame(msgs)

相关问题更多 >

编程相关推荐

热门问题

热门文章