[ https://issues.apache.org/jira/browse/BEAM-8335?focusedWorklogId=397149&page=com.atlassian.jira.plugin.system.issuetabpanels:worklog-tabpanel#worklog-397149 ]
ASF GitHub Bot logged work on BEAM-8335: ---------------------------------------- Author: ASF GitHub Bot Created on: 04/Mar/20 00:23 Start Date: 04/Mar/20 00:23 Worklog Time Spent: 10m Work Description: robertwb commented on pull request #11005: [BEAM-8335] Modify the StreamingCache to subclass the CacheManager URL: https://github.com/apache/beam/pull/11005#discussion_r387367517 ########## File path: sdks/python/apache_beam/runners/interactive/caching/streaming_cache.py ########## @@ -19,15 +19,298 @@ from __future__ import absolute_import +import itertools +import os +import shutil +import tempfile +import time + +import apache_beam as beam +from apache_beam.portability.api.beam_interactive_api_pb2 import TestStreamFileHeader +from apache_beam.portability.api.beam_interactive_api_pb2 import TestStreamFileRecord from apache_beam.portability.api.beam_runner_api_pb2 import TestStreamPayload +from apache_beam.runners.interactive.cache_manager import CacheManager +from apache_beam.runners.interactive.cache_manager import SafeFastPrimitivesCoder +from apache_beam.testing.test_stream import ReverseTestStream from apache_beam.utils import timestamp -class StreamingCache(object): +class StreamingCacheSink(beam.PTransform): + """A PTransform that writes TestStreamFile(Header|Records)s to file. + + This transform takes in an arbitrary element stream and writes the best-effort + list of TestStream events (as TestStreamFileRecords) to file. + + Note that this PTransform is assumed to be only run on a single machine where + the following assumptions are correct: elements come in ordered, no two + transforms are writing to the same file. This PTransform is assumed to only + run correctly with the DirectRunner. Review comment: Can you please file a JIRA to make this more general? (I don't see anything yet that would preclude writing something that works for all runners.) ---------------------------------------------------------------- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. For queries about this service, please contact Infrastructure at: us...@infra.apache.org Issue Time Tracking ------------------- Worklog Id: (was: 397149) Time Spent: 86h 10m (was: 86h) > Add streaming support to Interactive Beam > ----------------------------------------- > > Key: BEAM-8335 > URL: https://issues.apache.org/jira/browse/BEAM-8335 > Project: Beam > Issue Type: Improvement > Components: runner-py-interactive > Reporter: Sam Rohde > Assignee: Sam Rohde > Priority: Major > Time Spent: 86h 10m > Remaining Estimate: 0h > > This issue tracks the work items to introduce streaming support to the > Interactive Beam experience. This will allow users to: > * Write and run a streaming job in IPython > * Automatically cache records from unbounded sources > * Add a replay experience that replays all cached records to simulate the > original pipeline execution > * Add controls to play/pause/stop/step individual elements from the cached > records > * Add ability to inspect/visualize unbounded PCollections -- This message was sent by Atlassian Jira (v8.3.4#803005)