在Python中合并多行（文本格式）

Question

我正在通过一个网络接口获取日志，目前获取到的日志格式如下（下面有3个事件，都是以某个字符开头，以句点结束）。我的问题是，怎样才能逐行处理这些日志，并把它们连接起来，最终得到我想要的结果。

现在的输出

<attack_headlines version="1.0.1">
  <attack_headline>
    <site_id>1</site_id>
    <category>V2luZG93cyBEaXJlY3RvcmllcyBhbmQgRmlsZXM=</category>
    <subcategory>SUlTIEhlbHA=</subcategory>
    <client_ip>172.17.1.126</client_ip>
    <date>1363735940</date>
    <gmt_diff>0</gmt_diff>
    <reference_id>6D13-DE3D-9539-8980</reference_id>
  </attack_headline>
</attack_headlines>
<attack_headlines version="1.0.1">
  <attack_headline>
    <site_id>1</site_id>
    <category>V2luZG93cyBEaXJlY3RvcmllcyBhbmQgRmlsZXM=</category>
    <subcategory>SUlTIEhlbHA=</subcategory>
    <client_ip>172.17.1.136</client_ip>
    <date>1363735971</date>
    <gmt_diff>0</gmt_diff>
    <reference_id>6D13-DE3D-9539-8981</reference_id>
  </attack_headline>
</attack_headlines>
<attack_headlines version="1.0.1">
  <attack_headline>
    <site_id>1</site_id>
    <category>V2luZG93cyBEaXJlY3RvcmllcyBhbmQgRmlsZXM=</category>
    <subcategory>SUlTIEhlbHA=</subcategory>
    <client_ip>172.17.1.156</client_ip>
    <date>1363735975</date>
    <gmt_diff>0</gmt_diff>
    <reference_id>6D13-DE3D-9539-8982</reference_id>
  </attack_headline>
</attack_headlines>

期望的输出

<attack_headlines version="1.0.1"><attack_headline><site_id>1</site_id<category>V2luZG93cyBEaXJlY3RvcmllcyBhbmQgRmlsZXM=</category<subcategory>SUlTIEhlbHA=</subcategory><client_ip>172.17.1.156</client_ip<date>1363735975</date><gmt_diff>0</gmt_diff<reference_id>6D13-DE3D-9539-8982</reference_id></attack_headline</attack_headlines>

提前谢谢你们！

import json
import os
from suds.transport.https import WindowsHttpAuthenticated

class Helpers:
        def set_connection(self,conf):
                        #SUDS BUG FIXER(doctor)
                        protocol=conf['protocol']
                        hostname=conf['hostname']
                        port=conf['port']
                        path=conf['path']
                        file=conf['file']
                        u_name=conf['login']
                        passwrd=conf['password']
                        auth_type = conf['authType']

                        from suds.xsd.doctor import ImportDoctor, Import
                        from suds.client import Client

                        url = '{0}://{1}:{2}/{3}/{4}?wsdl'.format(protocol,
                        hostname,port, path, file)

                        imp = Import('http://schemas.xmlsoap.org/soap/encoding/')
                        d = ImportDoctor(imp)
                        if(auth_type == 'ntlm'):
                                ntlm = WindowsHttpAuthenticated(username=u_name, password=passwrd)
                                client = Client(url, transport=ntlm, doctor=d)
                        else:
                                client = Client(url, username=u_name, password=passwrd, doctor=d)
                        return client
        def read_from_file(self, filename):
                try:
                        fo = open(filename, "r")
                        try:
                                result = fo.read()
                        finally:
                                fo.close()
                                return result
                except IOError:
                        print "##Error opening/reading file {0}".format(filename)
                        exit(-1)


        def read_json(self,filename):
                string=self.read_from_file(filename)
                return json.loads(string)


        def get_recent_attacks(self, client):
            import time
            import base64
            from xml.dom.minidom import parseString
            epoch_time_now = int(time.time())
            epochtimeread = open('epoch_last', 'r')
            epoch_time_last_read = epochtimeread.read()
            epochtimeread.close()
            epoch_time_last = int(float(epoch_time_last_read))
            print client.service.get_recent_attacks("",epoch_time_last,epoch_time_now,1,"",15)

网络接口日志处理数据格式化行处理文本合并

在Python中合并多行（文本格式）

3 个回答

撰写回答