在Python数据流/apachebeam上启动CloudSQL代理

CUSTOM_COMMANDS = [ ['echo', 'Custom command worked!'], ['wget', 'https://dl.google.com/cloudsql/cloud_sql_proxy.linux.amd64', '-O', 'cloud_sql_proxy'], ['echo', 'Proxy downloaded'], ['chmod', '+x', 'cloud_sql_proxy']] class CustomCommands(setuptools.Command): """A setuptools Command class able to run arbitrary commands.""" def initialize_options(self): pass def finalize_options(self): pass def RunCustomCommand(self, command_list): print('Running command: %s' % command_list) logging.info("Running custom commands") p = subprocess.Popen( command_list, stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.STDOUT) # Can use communicate(input='y\n'.encode()) if the command run requires # some confirmation. stdout_data, _ = p.communicate() print('Command output: %s' % stdout_data) if p.returncode != 0: raise RuntimeError( 'Command %s failed: exit code: %s' % (command_list, p.returncode)) def run(self): for command in CUSTOM_COMMANDS: self.RunCustomCommand(command) subprocess.Popen(['./cloud_sql_proxy', '-instances=bi-test-1:europe-west1:test-animal=tcp:5432'])

更新日志文件2

作业日志（我在一段时间没有结巴后手动取消了作业）：

2018-06-08 (08:02:20) Autoscaling is enabled for job 2018-06-07_23_02_20-5917188751755240698. The number of workers will b... 2018-06-08 (08:02:20) Autoscaling was automatically enabled for job 2018-06-07_23_02_20-5917188751755240698. 2018-06-08 (08:02:24) Checking required Cloud APIs are enabled. 2018-06-08 (08:02:24) Checking permissions granted to controller Service Account. 2018-06-08 (08:02:25) Worker configuration: n1-standard-1 in europe-west1-b. 2018-06-08 (08:02:25) Expanding CoGroupByKey operations into optimizable parts. 2018-06-08 (08:02:25) Combiner lifting skipped for step Save new watermarks/Write/WriteImpl/GroupByKey: GroupByKey not fol... 2018-06-08 (08:02:25) Combiner lifting skipped for step Group watermarks: GroupByKey not followed by a combiner. 2018-06-08 (08:02:25) Expanding GroupByKey operations into optimizable parts. 2018-06-08 (08:02:26) Lifting ValueCombiningMappingFns into MergeBucketsMappingFns 2018-06-08 (08:02:26) Annotating graph with Autotuner information. 2018-06-08 (08:02:26) Fusing adjacent ParDo, Read, Write, and Flatten operations 2018-06-08 (08:02:26) Fusing consumer Get rows from CloudSQL tables into Begin pipeline with watermarks/Read 2018-06-08 (08:02:26) Fusing consumer Group watermarks/Write into Group watermarks/Reify 2018-06-08 (08:02:26) Fusing consumer Group watermarks/GroupByWindow into Group watermarks/Read 2018-06-08 (08:02:26) Fusing consumer Save new watermarks/Write/WriteImpl/WriteBundles/WriteBundles into Save new watermar... 2018-06-08 (08:02:26) Fusing consumer Save new watermarks/Write/WriteImpl/GroupByKey/GroupByWindow into Save new watermark... 2018-06-08 (08:02:26) Fusing consumer Save new watermarks/Write/WriteImpl/GroupByKey/Reify into Save new watermarks/Write/... 2018-06-08 (08:02:26) Fusing consumer Save new watermarks/Write/WriteImpl/GroupByKey/Write into Save new watermarks/Write/... 2018-06-08 (08:02:26) Fusing consumer Write to BQ into Get rows from CloudSQL tables 2018-06-08 (08:02:26) Fusing consumer Group watermarks/Reify into Write to BQ 2018-06-08 (08:02:26) Fusing consumer Save new watermarks/Write/WriteImpl/Map(<lambda at iobase.py:926>) into Convert dict... 2018-06-08 (08:02:26) Fusing consumer Save new watermarks/Write/WriteImpl/WindowInto(WindowIntoFn) into Save new watermark... 2018-06-08 (08:02:26) Fusing consumer Convert dictionary list to single dictionary and json into Remove "watermark" label 2018-06-08 (08:02:26) Fusing consumer Remove "watermark" label into Group watermarks/GroupByWindow 2018-06-08 (08:02:26) Fusing consumer Save new watermarks/Write/WriteImpl/InitializeWrite into Save new watermarks/Write/W... 2018-06-08 (08:02:26) Workflow config is missing a default resource spec. 2018-06-08 (08:02:26) Adding StepResource setup and teardown to workflow graph. 2018-06-08 (08:02:26) Adding workflow start and stop steps. 2018-06-08 (08:02:26) Assigning stage ids. 2018-06-08 (08:02:26) Executing wait step start25 2018-06-08 (08:02:26) Executing operation Save new watermarks/Write/WriteImpl/DoOnce/Read+Save new watermarks/Write/WriteI... 2018-06-08 (08:02:26) Executing operation Save new watermarks/Write/WriteImpl/GroupByKey/Create 2018-06-08 (08:02:26) Starting worker pool setup. 2018-06-08 (08:02:26) Executing operation Group watermarks/Create 2018-06-08 (08:02:26) Starting 1 workers in europe-west1-b... 2018-06-08 (08:02:27) Value "Group watermarks/Session" materialized. 2018-06-08 (08:02:27) Value "Save new watermarks/Write/WriteImpl/GroupByKey/Session" materialized. 2018-06-08 (08:02:27) Executing operation Begin pipeline with watermarks/Read+Get rows from CloudSQL tables+Write to BQ+Gr... 2018-06-08 (08:02:36) Autoscaling: Raised the number of workers to 0 based on the rate of progress in the currently runnin... 2018-06-08 (08:02:46) Autoscaling: Raised the number of workers to 1 based on the rate of progress in the currently runnin... 2018-06-08 (08:03:05) Workers have started successfully. 2018-06-08 (08:11:37) Cancel request is committed for workflow job: 2018-06-07_23_02_20-5917188751755240698. 2018-06-08 (08:11:38) Cleaning up. 2018-06-08 (08:11:38) Starting worker pool teardown. 2018-06-08 (08:11:38) Stopping worker pool... 2018-06-08 (08:12:30) Autoscaling: Reduced the number of workers to 0 based on the rate of progress in the currently runni...

堆栈跟踪：

No errors have been received in this time period.

1条回答

网友

1楼 · 发布于 2024-04-19 02:44:13

解决方案：

我终于找到了解决办法。我想到了通过CloudSQL实例的公共IP进行连接。为此，您需要允许从每个IP连接到CloudSQL实例：

转到GCP中CloudSQL实例的概述页面
单击Authorization选项卡
单击Add network并添加0.0.0.0/0（！！这将允许每个IP地址连接到您的实例！！）

为了给进程添加安全性，我使用了SSL密钥，并且只允许SSL连接到实例：

单击SSL选项卡
单击Create a new certificate为您的服务器创建一个SSL证书
单击Create a client certificate为您的客户机创建一个SSL证书
{cd7>拒绝所有连接尝试^无

之后，我将证书存储在Google云存储桶中并加载在数据流作业内连接之前，即：

import psycopg2
import psycopg2.extensions
import os
import stat
from google.cloud import storage

# Function to wait for open connection when processing parallel
def wait(conn):
    while 1:
        state = conn.poll()
        if state == psycopg2.extensions.POLL_OK:
            break
        elif state == psycopg2.extensions.POLL_WRITE:
            pass
            select.select([], [conn.fileno()], [])
        elif state == psycopg2.extensions.POLL_READ:
            pass
            select.select([conn.fileno()], [], [])
        else:
            raise psycopg2.OperationalError("poll() returned %s" % state)

# Function which returns a connection which can be used for queries
def connect_to_db(host, hostaddr, dbname, user, password, sslmode = 'verify-full'):

    # Get keys from GCS
    client = storage.Client()

    bucket = client.get_bucket(<YOUR_BUCKET_NAME>)

    bucket.get_blob('PATH_TO/server-ca.pem').download_to_filename('server-ca.pem')
    bucket.get_blob('PATH_TO/client-key.pem').download_to_filename('client-key.pem')
    os.chmod("client-key.pem", stat.S_IRWXU)
    bucket.get_blob('PATH_TO/client-cert.pem').download_to_filename('client-cert.pem')

    sslrootcert = 'server-ca.pem'
    sslkey = 'client-key.pem'
    sslcert = 'client-cert.pem'

    con = psycopg2.connect(
        host = host,
        hostaddr = hostaddr,
        dbname = dbname,
        user = user,
        password = password,
        sslmode=sslmode,
        sslrootcert = sslrootcert,
        sslcert = sslcert,
        sslkey = sslkey)
    return con

然后我在自定义ParDo中使用这些函数来执行查询。
最小示例：

^{pr2}$

管道的一部分可能如下所示：

# Current workaround to query all tables: 
# Create a dummy initiator PCollection with one element
init = p        |'Begin pipeline with initiator' >> beam.Create(['All tables initializer'])

tables = init   |'Get table names' >> beam.ParDo(ReadSQLTableNames(
                                                host = known_args.host,
                                                hostaddr = known_args.hostaddr,
                                                dbname = known_args.db_name,
                                                username = known_args.user,
                                                password = known_args.password))

我希望这个解决方案能帮助解决类似问题的其他人

更新-日志文件

更新日志文件2

更新：解决方法可以在我下面的答案中找到

解决方案：

相关问题更多 >

编程相关推荐

热门问题

热门文章