This is an automated email from the ASF dual-hosted git repository. github-merge-queue[bot] pushed a commit to branch gh-readonly-queue/main/pr-8544-dcf8217ff4487fdbb648c33c4e3f8b159d819a51 in repository https://gitbox.apache.org/repos/asf/texera.git
commit 3733b2a11948a9ed83db58864db980dc4d7d6bc6 Author: Ryan Zhang <[email protected]> AuthorDate: Wed Sep 16 21:35:45 2026 +0000 feat(notebook-migration): support migrating Python files into workflows (#8544) ### What changes were proposed in this PR? The migration tool accepted only a Jupyter notebook. It now accepts a Python file as well, and produces the same result: a workflow, a stored source file, and an operator mapping that drives highlighting in the Jupyter panel. The obstacle is that a notebook arrives already split into cells, and those cells are the join key for the mapping. A script has no such boundaries. Rather than guess at them with a heuristic splitter, this PR asks the model for them: the script is sent with a line number on each line, the model reports which line ranges became which UDF, and the cells are derived from that answer. The model is already deciding which code becomes which operator, so it is the right thing to ask. Everything after the model's reply is shared with the notebook path. Same mapping shape, same storage, same Jupyter upload, same highlight index. No backend, database, Helm, or dependency changes. Main pieces: * `script-segmentation.ts`, a pure module that reconciles the reported ranges and derives cells. Ranges can arrive reversed, overlapping, out of order, past the end of the file, or not at all. Every case is reconciled rather than rejected, because the ranges come back alongside a workflow that already cost a full conversion. It cuts the file at the union of all reported boundaries, which makes overlaps and gaps fall out of one mechanism: an overlap becomes a shared cell mapped to both UDFs, and a gap becomes a cell no UDF claims, so no source is ever lost. * Script variants of the conversion prompt and of the worked example in the documentation prelude. The existing worked example is written in `# START CELL1` form and seeded as a system message, so a script conversion needed its own, or the model would answer in cell ids for an input that has none. `migration-prompts.ts` is 175 insertions and 1 deletion. All but five of those lines are the new script constants. The notebook prompts take two changes, both fixes to pre-existing bugs found during review: `MAPPING_PROMPT`'s example mapped only four of the five UDFs its worked example defines, which taught the model that leaving a UDF unmapped is acceptable, and `WORKFLOW_PROMPT` carried a malformed sentence ("every distinct UDF that uses that constructs an object of that class") that left the class duplication rule ambiguous. Nothing else in the notebook path changed. * `convertScriptToWorkflow` on the LLM client, plus `parseScriptFile` and `sendScriptToAIGenerateWorkflow` on the service. * A second tab in the AI generate modal, and extension dispatch in the dashboard entry point. Two behavior-neutral refactors support this: workflow assembly and mapping inversion were extracted out of `convertNotebookToWorkflow`, and the LLM session lifecycle (initialize, verify, close in a `finally`) was extracted so both paths share it and neither can leak a session. One deliberate behavior change rides along. The shared mapping inversion now skips a UDF whose cell list is not an array, with a warning, rather than throwing. A model that answers with a bare string there previously discarded an entire conversion after two expensive calls; it now degrades the mapping instead, which is what the script path already does for a malformed range. #### Demo video (waiting was cut out to keep demo short) https://github.com/user-attachments/assets/2e248258-4520-4117-b67b-d1e10ed5bd5c ### Any related issues, documentation, discussions? Closes #8007 Parent-issue #4301. The choice to keep the mapping and Jupyter panel for scripts, rather than returning only a workflow, was discussed in https://github.com/apache/texera/discussions/8489. The panel and toolbar copy still says "Jupyter Notebook" where a script user would expect something general; that is being handled in #8543 so this PR stays behavioral. ### How was this PR tested? Unit tests, all new unless noted: * `script-segmentation.spec.ts`: 32 tests over well-formed input, gaps, overlaps, malformed and out of bounds ranges, the accepted range shapes, degenerate input, and source fidelity. One test asserts that every line with content lands in exactly one cell, in order, which is the property that protects users from losing code. A blank run isolated between two reported ranges is dropped on purpose, so the cells do not reassemble byte for byte. * `migration-llm.spec.ts`: 31 to 42. Covers line numbering, prelude selection, cell derivation, the notebook it returns, degrading to one unmapped cell when no usable ranges come back, and the non-array mapping entry described above. * `notebook-migration.service.spec.ts`: 39 to 48. Lifecycle and `parseScriptFile`. * `notebook-import-modal.component.spec.ts`: 19 to 27. Tab switching, per tab upload targets, and submission. * `user-workflow.component.spec.ts`: 76 to 81. Extension dispatch, and that a `.py` stores the derived notebook and never reaches the notebook path. Full frontend suite passes (218 files, 5991 tests), along with `yarn format:ci` and a production build. Manually tested end to end against a live model: uploaded a `.py`, confirmed the generated workflow, confirmed the derived notebook opens in the Jupyter panel, and confirmed clicking an operator highlights the cell its code came from. ### Was this PR authored or co-authored using generative AI tooling? Generated-by: Claude Code (Claude Opus 5) --- .../user-workflow/user-workflow.component.html | 4 +- .../user-workflow/user-workflow.component.spec.ts | 102 +++++- .../user/user-workflow/user-workflow.component.ts | 102 ++++-- .../notebook-import-modal.component.html | 173 ++++++---- .../notebook-import-modal.component.scss | 29 +- .../notebook-import-modal.component.spec.ts | 149 ++++++++- .../notebook-import-modal.component.ts | 34 +- .../notebook-migration/migration-llm.spec.ts | 174 ++++++++++ .../service/notebook-migration/migration-llm.ts | 203 +++++++++--- .../notebook-migration/migration-prompts.ts | 176 ++++++++++- .../notebook-migration.service.spec.ts | 94 ++++++ .../notebook-migration.service.ts | 85 ++++- .../notebook-migration/script-segmentation.spec.ts | 350 +++++++++++++++++++++ .../notebook-migration/script-segmentation.ts | 191 +++++++++++ .../python_file_diagram.png | Bin 0 -> 71396 bytes 15 files changed, 1723 insertions(+), 143 deletions(-) diff --git a/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.html b/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.html index f4e949c4fa..7adcd92a4c 100644 --- a/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.html +++ b/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.html @@ -53,8 +53,8 @@ *ngIf="pythonNotebookMigrationEnabled" nz-button (click)="openAiGenerateModal()" - title="AI generate a workflow from a Python notebook" - nz-tooltip="AI generate a workflow from a Python notebook" + title="AI generate a workflow from source code" + nz-tooltip="AI generate a workflow from source code" nzTooltipPlacement="bottom" type="button"> <i diff --git a/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.spec.ts b/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.spec.ts index 0986d60b17..764b4bdc11 100644 --- a/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.spec.ts +++ b/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.spec.ts @@ -317,7 +317,7 @@ describe("SavedWorkflowSectionComponent", () => { describe("AI generate workflow (dashboard entry point)", () => { const ipynbFile = { name: "analysis.ipynb" } as NzUploadFile; - const AI_BUTTON_SELECTOR = 'button[title="AI generate a workflow from a Python notebook"]'; + const AI_BUTTON_SELECTOR = 'button[title="AI generate a workflow from source code"]'; // Opens the modal and returns the requestImport callback the component handed to it; calling // it runs the full generation (true => generation succeeded and navigated, false => stay open). @@ -353,7 +353,9 @@ describe("SavedWorkflowSectionComponent", () => { expect(createSpy).toHaveBeenCalledTimes(1); const config = createSpy.mock.calls[0][0] as ModalOptions; expect(config.nzContent).toBe(NotebookImportModalComponent); + expect(config.nzTitle).toBe("AI Generate Workflow from Source Code"); expect(config.nzFooter).toBeNull(); + expect(config.nzBodyStyle).toEqual({ paddingTop: "4px" }); expect(typeof (config.nzData as { requestImport: unknown }).requestImport).toBe("function"); }); @@ -399,15 +401,105 @@ describe("SavedWorkflowSectionComponent", () => { expect(proceed).toBe(true); }); - it("rejects a non-ipynb file: errors, resolves false, and generates nothing", async () => { - const parseSpy = vi.spyOn(TestBed.inject(NotebookMigrationService), "parseAndTagNotebook"); + it("rejects an unsupported extension: errors, resolves false, and generates nothing", async () => { + const migration = TestBed.inject(NotebookMigrationService); + const notebookSpy = vi.spyOn(migration, "parseAndTagNotebook"); + const scriptSpy = vi.spyOn(migration, "parseScriptFile"); const errorSpy = vi.spyOn(TestBed.inject(NotificationService), "error").mockImplementation(() => {}); const proceed = await getRequestImport()({ name: "data.txt" } as NzUploadFile, "gpt-4"); expect(proceed).toBe(false); - expect(errorSpy).toHaveBeenCalledWith("Please upload a valid Jupyter Notebook (.ipynb) file."); - expect(parseSpy).not.toHaveBeenCalled(); + expect(errorSpy).toHaveBeenCalledWith("Please upload a Jupyter Notebook (.ipynb) or a Python (.py) file."); + expect(notebookSpy).not.toHaveBeenCalled(); + expect(scriptSpy).not.toHaveBeenCalled(); + }); + + // A .py takes the other branch: read as text, converted by the script method, and the + // notebook it stores is the one the LLM derived rather than an uploaded file. + describe("Python file input", () => { + const pyFile = { name: "analysis.py" } as NzUploadFile; + const derivedNotebook = { cells: [{ cell_type: "code", metadata: { uuid: "u1" }, source: "x = 1" }] }; + + function mockScriptGenerationSuccess(wid = 42) { + const migration = TestBed.inject(NotebookMigrationService); + vi.spyOn(migration, "parseScriptFile").mockResolvedValue("x = 1\n"); + const convertSpy = vi.spyOn(migration, "sendScriptToAIGenerateWorkflow").mockResolvedValue({ + workflowContent: { operators: [] }, + mappingContent: { operator_to_cell: {}, cell_to_operator: {} }, + notebook: derivedNotebook, + } as any); + const storeSpy = vi.spyOn(migration, "storeNotebookAndMapping").mockReturnValue(of({ success: true }) as any); + const persist = TestBed.inject(WorkflowPersistService) as any; + persist.createWorkflow = vi.fn().mockReturnValue(of({ workflow: { wid } })); + return { convertSpy, storeSpy, persist }; + } + + it("converts the script, stores the derived notebook, navigates, and resolves true", async () => { + const { convertSpy, storeSpy, persist } = mockScriptGenerationSuccess(42); + const navigateSpy = vi.spyOn(TestBed.inject(Router), "navigate").mockResolvedValue(true); + + const proceed = await getRequestImport()(pyFile, "gpt-4"); + + expect(convertSpy).toHaveBeenCalledWith("x = 1\n", "gpt-4"); + expect(persist.createWorkflow.mock.calls[0][1]).toBe("analysis_GENERATED_BY_LLM"); + // The stored notebook is the derived one; nothing else could have supplied it. + expect(storeSpy).toHaveBeenCalledWith(42, expect.anything(), derivedNotebook); + expect(navigateSpy).toHaveBeenCalledWith([USER_WORKSPACE, 42], { queryParams: { autolayout: 1 } }); + expect(proceed).toBe(true); + }); + + it("never reaches the notebook path for a .py", async () => { + mockScriptGenerationSuccess(); + vi.spyOn(TestBed.inject(Router), "navigate").mockResolvedValue(true); + const migration = TestBed.inject(NotebookMigrationService); + const notebookParse = vi.spyOn(migration, "parseAndTagNotebook"); + const notebookConvert = vi.spyOn(migration, "sendToAIGenerateWorkflow"); + + await getRequestImport()(pyFile, "gpt-4"); + + expect(notebookParse).not.toHaveBeenCalled(); + expect(notebookConvert).not.toHaveBeenCalled(); + }); + + it("reports an unreadable or empty file and resolves false without calling the LLM", async () => { + const migration = TestBed.inject(NotebookMigrationService); + vi.spyOn(migration, "parseScriptFile").mockRejectedValue(new Error("The Python file is empty.")); + const convertSpy = vi.spyOn(migration, "sendScriptToAIGenerateWorkflow"); + const errorSpy = vi.spyOn(TestBed.inject(NotificationService), "error").mockImplementation(() => {}); + + const proceed = await getRequestImport()(pyFile, "gpt-4"); + + expect(proceed).toBe(false); + expect(errorSpy).toHaveBeenCalledWith( + "Failed to read the Python file. Please upload a valid, non-empty .py file." + ); + expect(convertSpy).not.toHaveBeenCalled(); + }); + + it("names the script in the timeout message so the advice matches the input", async () => { + const migration = TestBed.inject(NotebookMigrationService); + vi.spyOn(migration, "parseScriptFile").mockResolvedValue("x = 1\n"); + vi.spyOn(migration, "sendScriptToAIGenerateWorkflow").mockRejectedValue(new LlmRequestTimeoutError(10)); + const errorSpy = vi.spyOn(TestBed.inject(NotificationService), "error").mockImplementation(() => {}); + + const proceed = await getRequestImport()(pyFile, "gpt-4"); + + expect(proceed).toBe(false); + expect(errorSpy).toHaveBeenCalledWith(expect.stringContaining("simplify the script")); + }); + + it("reports a generation failure and resolves false", async () => { + const migration = TestBed.inject(NotebookMigrationService); + vi.spyOn(migration, "parseScriptFile").mockResolvedValue("x = 1\n"); + vi.spyOn(migration, "sendScriptToAIGenerateWorkflow").mockRejectedValue(new Error("LLM down")); + const errorSpy = vi.spyOn(TestBed.inject(NotificationService), "error").mockImplementation(() => {}); + + const proceed = await getRequestImport()(pyFile, "gpt-4"); + + expect(proceed).toBe(false); + expect(errorSpy).toHaveBeenCalledWith("Error while communicating with the LLM, check console for details."); + }); }); it("reports a parse failure and resolves false without calling the LLM", async () => { diff --git a/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.ts b/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.ts index 5aae5f7943..8cb9423d7f 100644 --- a/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.ts +++ b/frontend/src/app/dashboard/component/user/user-workflow/user-workflow.component.ts @@ -65,6 +65,17 @@ import { FiltersInstructionsComponent } from "../filters-instructions/filters-in import { NzSelectComponent } from "ng-zorro-antd/select"; import { FormsModule } from "@angular/forms"; +/** + * What a conversion yields, whichever input produced it. `notebook` is the uploaded one for an + * .ipynb and the LLM-derived one for a .py; either way it is what gets stored and shown in the + * Jupyter panel. + */ +interface GeneratedWorkflow { + workflowContent: WorkflowContent; + mappingContent: MappingContent; + notebook: Notebook; +} + /** * Saved-workflow-section component contains information and functionality * of the saved workflows section: the list of workflows the user owns or has access to @@ -280,54 +291,103 @@ export class UserWorkflowComponent implements AfterViewInit, OnDestroy { return this.config.env.pythonNotebookMigrationEnabled; } - /** Open the AI-generate import modal, wiring its submit to generateWorkflowFromNotebook. */ + /** Open the AI-generate import modal, wiring its submit to generateWorkflowFromFile. */ public openAiGenerateModal(): void { this.modalService.create<NotebookImportModalComponent, NotebookImportModalData>({ - nzTitle: "AI Generate Workflow from Python Notebook", + nzTitle: "AI Generate Workflow from Source Code", nzContent: NotebookImportModalComponent, nzWidth: 700, nzFooter: null, nzCentered: true, + nzBodyStyle: { paddingTop: "4px" }, nzData: { - requestImport: (file, model) => this.generateWorkflowFromNotebook(file, model), + requestImport: (file, model) => this.generateWorkflowFromFile(file, model), }, }); } /** - * Parse the notebook, generate a workflow via the LLM, save it, store the cell mapping, and open it. - * Resolves true on success (modal closes), false to keep the modal open on a bad file or a failure. + * Generate a workflow from an uploaded notebook or Python file, save it, store the mapping, + * and open it. Resolves true on success (modal closes), false to keep the modal open on a bad + * file or a failure. */ - private async generateWorkflowFromNotebook(file: NzUploadFile, model: string): Promise<boolean> { + private async generateWorkflowFromFile(file: NzUploadFile, model: string): Promise<boolean> { const fileExtension = file.name.split(".").pop()?.toLowerCase(); - if (fileExtension !== "ipynb") { - this.notificationService.error("Please upload a valid Jupyter Notebook (.ipynb) file."); + if (fileExtension !== "ipynb" && fileExtension !== "py") { + this.notificationService.error("Please upload a Jupyter Notebook (.ipynb) or a Python (.py) file."); return false; } + + const generated = + fileExtension === "ipynb" + ? await this.generateFromNotebook(file, model) + : await this.generateFromScript(file, model); + // Null means the step already told the user what went wrong. + if (!generated) { + return false; + } + + return this.saveAndOpenGenerated(file, generated); + } + + /** Read and convert an .ipynb. Null after reporting a read or generation failure. */ + private async generateFromNotebook(file: NzUploadFile, model: string): Promise<GeneratedWorkflow | null> { let notebook: Notebook; try { notebook = await this.notebookMigrationService.parseAndTagNotebook(file as unknown as File); } catch (error) { this.notificationService.error("Failed to read the notebook file. Please upload a valid .ipynb file."); console.error("Notebook parse failed:", error); - return false; + return null; } - let generated: { workflowContent: WorkflowContent; mappingContent: MappingContent }; try { - generated = await this.notebookMigrationService.sendToAIGenerateWorkflow(notebook, model); + const generated = await this.notebookMigrationService.sendToAIGenerateWorkflow(notebook, model); + // The uploaded notebook is what gets stored, so it rides along with the generated pair. + return { ...generated, notebook }; } catch (error) { - if (error instanceof LlmRequestTimeoutError) { - this.notificationService.error( - `Generation timed out after ${error.minutes} minutes. Try again, choose a faster model, or simplify the notebook.` - ); - } else { - this.notificationService.error("Error while communicating with the LLM, check console for details."); - } - console.error("LLM generation failed:", error); - return false; + this.reportGenerationFailure(error, "notebook"); + return null; + } + } + + /** Read and convert a .py. The notebook comes back derived, since the upload had no cells. */ + private async generateFromScript(file: NzUploadFile, model: string): Promise<GeneratedWorkflow | null> { + let scriptSource: string; + try { + scriptSource = await this.notebookMigrationService.parseScriptFile(file as unknown as File); + } catch (error) { + this.notificationService.error("Failed to read the Python file. Please upload a valid, non-empty .py file."); + console.error("Python file read failed:", error); + return null; + } + + try { + return await this.notebookMigrationService.sendScriptToAIGenerateWorkflow(scriptSource, model); + } catch (error) { + this.reportGenerationFailure(error, "script"); + return null; } + } + + // A timeout is worth telling apart from a transport error: the user can act on it by picking + // a faster model or trimming the input. + private reportGenerationFailure(error: unknown, input: "notebook" | "script"): void { + if (error instanceof LlmRequestTimeoutError) { + this.notificationService.error( + `Generation timed out after ${error.minutes} minutes. Try again, choose a faster model, or simplify the ${input}.` + ); + } else { + this.notificationService.error("Error while communicating with the LLM, check console for details."); + } + console.error("LLM generation failed:", error); + } + /** + * Persist the generated workflow, attach its notebook and mapping, then open it. Shared by both + * inputs: once a conversion has produced a workflow and a notebook, nothing downstream differs. + */ + private async saveAndOpenGenerated(file: NzUploadFile, generated: GeneratedWorkflow): Promise<boolean> { // Commit point: persisting captures the expensive LLM result. On failure nothing was created, // so returning false to let the user retry is safe. let wid: number; @@ -351,7 +411,7 @@ export class UserWorkflowComponent implements AfterViewInit, OnDestroy { // Best-effort follow-ups: never discard the created workflow, so log/warn and still open it. try { await firstValueFrom( - this.notebookMigrationService.storeNotebookAndMapping(wid, generated.mappingContent, notebook) + this.notebookMigrationService.storeNotebookAndMapping(wid, generated.mappingContent, generated.notebook) ); } catch (error) { this.notificationService.warning( diff --git a/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.html b/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.html index 814b92069c..5fc6e8f5c6 100644 --- a/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.html +++ b/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.html @@ -17,69 +17,130 @@ under the License. --> -<div class="import-modal-diagram"> - <img - ngSrc="assets/notebook_migration_tool/tool_popup_diagram.png" - alt="Notebook to Workflow" - width="1132" - height="290" /> -</div> - <div class="import-modal-content"> <form class="import-modal-form" [formGroup]="importForm" [attr.inert]="isSubmitting ? '' : null" nz-form> - <nz-form-item> - <p class="import-modal-text"> - This tool converts a Python Jupyter Notebook into a Texera workflow using LLM capabilities. After you submit a - notebook, the LLM service generates a corresponding Texera workflow. The conversion time depends on the - notebook's complexity and can take 1-5 minutes. Once generation finishes, you are taken to the new workflow, - which opens with: - </p> - <ol class="import-modal-list"> - <li> - The generated workflow ready to use (Note: you will still need to upload the dataset and connect it to the - workflow). - </li> - <li>A floating Jupyter window containing the uploaded notebook for reference.</li> - </ol> - <p class="import-modal-text"> - Generation runs here after you submit. Please keep this window open while you wait. - </p> - </nz-form-item> + <!-- One upload row shared by both tabs. They differ only in the accepted extension and + their labels, so the markup lives here once and each tab passes its own context. --> + <ng-template + #uploadRow + let-accept="accept" + let-label="label" + let-action="action"> + <nz-form-item> + <nz-form-label [nzNoColon]="true"> + <span class="import-modal-label">{{ label }}</span> + </nz-form-label> + <nz-form-control> + <div class="import-modal-upload-row"> + <nz-upload + [nzAccept]="accept" + [nzBeforeUpload]="beforeUpload" + [nzShowUploadList]="false"> + <button + nz-button + type="button" + [title]="action" + [attr.aria-label]="action"> + <i + nz-icon + nzType="upload"></i> + </button> + </nz-upload> - <nz-form-item> - <nz-form-label [nzNoColon]="true"> - <span class="import-modal-label"> Upload Python Jupyter Notebook </span> - </nz-form-label> - <nz-form-control> - <div class="import-modal-upload-row"> - <nz-upload - nzAccept=".ipynb" - [nzBeforeUpload]="beforeUpload" - [nzShowUploadList]="false"> - <button - nz-button - type="button" - title="Upload notebook" - aria-label="Upload notebook"> - <i - nz-icon - nzType="upload"></i> - </button> - </nz-upload> + <span + *ngIf="importForm.get('file')?.value?.name as fileName" + class="import-modal-selected-file" + [title]="fileName"> + Selected file: {{ fileName }} + </span> + </div> + </nz-form-control> + </nz-form-item> + </ng-template> - <span - *ngIf="importForm.get('file')?.value?.name as fileName" - class="import-modal-selected-file" - [title]="fileName"> - Selected file: {{ fileName }} - </span> - </div> - </nz-form-control> - </nz-form-item> + <nz-tabs + [nzSelectedIndex]="selectedTabIndex" + [nzAnimated]="tabAnimation" + (nzSelectedIndexChange)="onTabChange($event)"> + <nz-tab nzTitle="Jupyter Notebook"> + <ng-template nz-tab> + <div class="import-modal-diagram"> + <img + ngSrc="assets/notebook_migration_tool/tool_popup_diagram.png" + alt="Notebook to Workflow" + width="1132" + height="290" /> + </div> + + <nz-form-item> + <p class="import-modal-text"> + This tool converts a Python Jupyter Notebook into a Texera workflow using LLM capabilities. After you + submit a notebook, the LLM service generates a corresponding Texera workflow. The conversion time depends + on the notebook's complexity and can take 1-5 minutes. Once generation finishes, you are taken to the new + workflow, which opens with: + </p> + <ol class="import-modal-list"> + <li> + The generated workflow ready to use (Note: you will still need to upload the dataset and connect it to + the workflow). + </li> + <li>A floating Jupyter window containing the uploaded notebook for reference.</li> + </ol> + <p class="import-modal-text"> + Generation runs here after you submit. Please keep this window open while you wait. + </p> + </nz-form-item> + + <ng-container + *ngTemplateOutlet=" + uploadRow; + context: { accept: '.ipynb', label: 'Upload Python Jupyter Notebook', action: 'Upload notebook' } + "></ng-container> + </ng-template> + </nz-tab> + + <nz-tab nzTitle="Python File"> + <ng-template nz-tab> + <div class="import-modal-diagram"> + <img + ngSrc="assets/notebook_migration_tool/python_file_diagram.png" + alt="Python File to Workflow" + width="1348" + height="346" /> + </div> + + <nz-form-item> + <p class="import-modal-text"> + This tool converts a Python file into a Texera workflow using LLM capabilities. After you submit a script, + the LLM service generates a corresponding Texera workflow. A script has no notebook cells, so the LLM also + divides it into sections and records which section became which operator. The conversion time depends on + the script's complexity and can take 1-5 minutes. Once generation finishes, you are taken to the new + workflow, which opens with: + </p> + <ol class="import-modal-list"> + <li> + The generated workflow ready to use (Note: you will still need to upload the dataset and connect it to + the workflow). + </li> + <li>A floating Jupyter window containing your script, divided into cells, for reference.</li> + </ol> + <p class="import-modal-text"> + Generation runs here after you submit. Please keep this window open while you wait. + </p> + </nz-form-item> + + <ng-container + *ngTemplateOutlet=" + uploadRow; + context: { accept: '.py', label: 'Upload Python File', action: 'Upload Python file' } + "></ng-container> + </ng-template> + </nz-tab> + </nz-tabs> <nz-form-item> <nz-form-label [nzNoColon]="true"> @@ -147,7 +208,7 @@ [nzSize]="'large'"></nz-spin> <p class="import-modal-loading-title">Generating your workflow</p> <p class="import-modal-loading-elapsed">Elapsed time: {{ formattedElapsedTime }}</p> - <p class="import-modal-loading-text">This can take 1-5 minutes depending on the notebook's complexity.</p> + <p class="import-modal-loading-text">This can take 1-5 minutes depending on the code's complexity.</p> <p class="import-modal-loading-text"> Please keep this window open. You will be taken to the new workflow when it is ready. </p> diff --git a/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.scss b/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.scss index ed2b53d725..862f9a4f37 100644 --- a/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.scss +++ b/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.scss @@ -18,9 +18,23 @@ */ .import-modal { - // Tighten the default form-item spacing so the modal fits without scrolling. - &-form nz-form-item { - margin-bottom: 12px; + &-form { + // Takes up the slack under the fixed-height body so the footer stays at the bottom, + // and scrolls only if a tab's content ever outgrows the box. + flex: 1 1 auto; + min-height: 0; + overflow-y: auto; + + // Tighten the default form-item spacing so the modal fits without scrolling. + nz-form-item { + margin-bottom: 12px; + } + + // Tabs sit flush with the modal body; the default pane padding would double the + // spacing already provided by the diagram and form items inside each tab. + ::ng-deep .ant-tabs-tabpane { + padding-top: 4px; + } } &-diagram { @@ -79,9 +93,15 @@ width: 50%; } - // Positioning context for the loading overlay, so covering the form does not resize the modal. + // Positioning context for the loading overlay. The fixed height is sized to the taller tab so + // switching tabs does not resize the modal; max-height clamps it on short viewports, where the + // form's own overflow-y takes over. &-content { position: relative; + height: 620px; + max-height: calc(100vh - 140px); + display: flex; + flex-direction: column; } &-loading { @@ -122,6 +142,7 @@ &-footer { display: flex; + align-items: center; justify-content: flex-end; gap: 8px; margin-top: 12px; diff --git a/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.spec.ts b/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.spec.ts index 2e46f2ee65..ee31753219 100644 --- a/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.spec.ts +++ b/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.spec.ts @@ -222,9 +222,16 @@ describe("NotebookImportModalComponent", () => { } }); - it("has a visibilitychange handler that is safe to call (repaint is driven by the zone event)", async () => { + it("wires document visibilitychange to the host listener that repaints the stopwatch", async () => { await createWith(of([{ name: "gpt-4" }])); - expect(() => component.onVisibilityChange()).not.toThrow(); + // Spying after creation still observes the call: the generated host handler resolves + // ctx.onVisibilityChange at dispatch time. The method body is empty on purpose, so the + // wiring is the only thing worth asserting. + const handler = vi.spyOn(component, "onVisibilityChange"); + + document.dispatchEvent(new Event("visibilitychange")); + + expect(handler).toHaveBeenCalled(); }); it("clears the stopwatch interval on destroy", async () => { @@ -286,4 +293,142 @@ describe("NotebookImportModalComponent", () => { expect(modalRef.close).toHaveBeenCalledWith(); expect(requestImport).not.toHaveBeenCalled(); }); + + it("swaps the diagram, description and upload target when the Python tab is selected", async () => { + await createWith(of([{ name: "gpt-4" }])); + const root = fixture.nativeElement as HTMLElement; + // Once rendered, a tab pane stays in the DOM and is only hidden, so assert on the + // active pane rather than the whole modal. + const pane = () => root.querySelector(".ant-tabs-tabpane-active") as HTMLElement; + const accept = () => pane().querySelector("input[type='file']")?.getAttribute("accept"); + + // The src is asserted, not just the alt: a wrong path would otherwise pass here and 404 in + // the browser. ngSrc rewrites the attribute, so match on a suffix rather than the whole value. + const diagramSrc = () => pane().querySelector("img")?.getAttribute("src") ?? ""; + + expect(pane().querySelector("img[alt='Notebook to Workflow']")).not.toBeNull(); + expect(diagramSrc()).toContain("assets/notebook_migration_tool/tool_popup_diagram.png"); + expect(accept()).toBe(".ipynb"); + + component.onTabChange(1); + fixture.detectChanges(); + + expect(pane().querySelector("img[alt='Notebook to Workflow']")).toBeNull(); + expect(pane().querySelector("img[alt='Python File to Workflow']")).not.toBeNull(); + expect(diagramSrc()).toContain("assets/notebook_migration_tool/python_file_diagram.png"); + expect(accept()).toBe(".py"); + expect(pane().textContent).toContain("Upload Python File"); + }); + + it("drops the selected file when switching tabs so it cannot be staged under the wrong extension", async () => { + await createWith(of([{ name: "gpt-4" }])); + component.importForm.setValue({ file: { name: "demo.ipynb" }, model: "gpt-4" }); + + component.onTabChange(1); + fixture.detectChanges(); + + expect(component.importForm.get("file")?.value).toBeNull(); + // The model applies to either input, so it survives the switch. + expect(component.importForm.get("model")?.value).toBe("gpt-4"); + expect((fixture.nativeElement as HTMLElement).textContent).not.toContain("Selected file:"); + }); + + it("enables Submit on the Python tab once a file and model are chosen", async () => { + await createWith(of([{ name: "gpt-4" }])); + const root = fixture.nativeElement as HTMLElement; + const submit = () => root.querySelector(".import-modal-footer button[nzType='primary']") as HTMLButtonElement; + + component.onTabChange(1); + fixture.detectChanges(); + // Switching tabs clears the file, so the form is incomplete until one is staged. + expect(submit().disabled).toBe(true); + + component.importForm.setValue({ file: { name: "demo.py" }, model: "gpt-4" }); + fixture.detectChanges(); + + expect(submit().disabled).toBe(false); + }); + + it("submits a Python file to the opener the same way it submits a notebook", async () => { + requestImport.mockResolvedValue(true); + await createWith(of([{ name: "gpt-4" }])); + const file = { name: "demo.py" } as NzUploadFile; + component.onTabChange(1); + component.importForm.setValue({ file, model: "gpt-4" }); + + await component.onSubmit(); + + expect(requestImport).toHaveBeenCalledWith(file, "gpt-4"); + expect(modalRef.close).toHaveBeenCalledWith(); + }); + + it("the footer Cancel button closes the modal", async () => { + await createWith(of([{ name: "gpt-4" }])); + const cancel = (fixture.nativeElement as HTMLElement).querySelector( + ".import-modal-footer button:not([nzType='primary'])" + ) as HTMLButtonElement; + + // Clicked rather than calling onCancel(), so the template's click binding is exercised too. + cancel.click(); + + expect(modalRef.close).toHaveBeenCalledWith(); + expect(requestImport).not.toHaveBeenCalled(); + }); + + it("the footer Submit button runs the import with the staged file and model", async () => { + requestImport.mockResolvedValue(true); + await createWith(of([{ name: "gpt-4" }])); + const file = { name: "demo.ipynb" } as NzUploadFile; + component.importForm.setValue({ file, model: "gpt-4" }); + fixture.detectChanges(); + + // Clicked rather than calling onSubmit(), so the template's click binding is exercised too. + ( + (fixture.nativeElement as HTMLElement).querySelector( + ".import-modal-footer button[nzType='primary']" + ) as HTMLButtonElement + ).click(); + await fixture.whenStable(); + + expect(requestImport).toHaveBeenCalledWith(file, "gpt-4"); + expect(modalRef.close).toHaveBeenCalledWith(); + }); + + it("selects the Python tab from a click on its header, not just a direct call", async () => { + await createWith(of([{ name: "gpt-4" }])); + const headers = (fixture.nativeElement as HTMLElement).querySelectorAll(".ant-tabs-tab"); + expect(headers.length).toBe(2); + + (headers[1] as HTMLElement).click(); + // nz-tabs emits nzSelectedIndexChange from a microtask inside ngAfterContentChecked, + // so the handler has not run until the queue drains. + fixture.detectChanges(); + await fixture.whenStable(); + fixture.detectChanges(); + + expect(component.selectedTabIndex).toBe(1); + expect( + (fixture.nativeElement as HTMLElement) + .querySelector(".ant-tabs-tabpane-active input[type='file']") + ?.getAttribute("accept") + ).toBe(".py"); + }); + + it("restores the notebook upload target when switching back", async () => { + await createWith(of([{ name: "gpt-4" }])); + const root = fixture.nativeElement as HTMLElement; + const pane = () => root.querySelector(".ant-tabs-tabpane-active") as HTMLElement; + const submit = () => root.querySelector(".import-modal-footer button[nzType='primary']") as HTMLButtonElement; + + component.onTabChange(1); + fixture.detectChanges(); + expect(pane().querySelector("input[type='file']")?.getAttribute("accept")).toBe(".py"); + + component.onTabChange(0); + component.importForm.setValue({ file: { name: "demo.ipynb" }, model: "gpt-4" }); + fixture.detectChanges(); + + expect(pane().querySelector("input[type='file']")?.getAttribute("accept")).toBe(".ipynb"); + expect(submit().disabled).toBe(false); + }); }); diff --git a/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.ts b/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.ts index 7c9ad86cd6..5d2c6895ae 100644 --- a/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.ts +++ b/frontend/src/app/workspace/component/notebook-import-modal/notebook-import-modal.component.ts @@ -22,12 +22,13 @@ import { FormBuilder, FormGroup, Validators, ReactiveFormsModule } from "@angula import { NZ_MODAL_DATA, NzModalRef } from "ng-zorro-antd/modal"; import { NzUploadComponent, NzUploadFile } from "ng-zorro-antd/upload"; import { Observable } from "rxjs"; -import { AsyncPipe, NgIf, NgFor, NgOptimizedImage } from "@angular/common"; +import { AsyncPipe, NgIf, NgFor, NgOptimizedImage, NgTemplateOutlet } from "@angular/common"; import { NzFormModule } from "ng-zorro-antd/form"; import { NzSelectModule } from "ng-zorro-antd/select"; import { NzSpinComponent } from "ng-zorro-antd/spin"; import { NzButtonComponent } from "ng-zorro-antd/button"; import { NzIconDirective } from "ng-zorro-antd/icon"; +import { NzTabsComponent, NzTabComponent, NzTabDirective } from "ng-zorro-antd/tabs"; import { NotebookMigrationService } from "../../service/notebook-migration/notebook-migration.service"; // Passed in via nzData. requestImport resolves true to close the modal, false to keep it open @@ -37,8 +38,12 @@ export interface NotebookImportModalData { } /** - * The "AI Generate Workflow from Python Notebook" modal body: the upload form and model dropdown. - * On Submit it hands the file and model to requestImport and shows a loading state until it resolves. + * The "AI Generate Workflow from Source Code" modal body: a tab per accepted input kind + * (Jupyter notebook, Python script), each with its own diagram, description and upload + * control, over a shared model dropdown and footer. + * + * On Submit it hands the file and model to requestImport and shows a loading state until + * it resolves; the opener dispatches on the uploaded file's extension. */ @Component({ selector: "texera-notebook-import-modal", @@ -49,6 +54,7 @@ export interface NotebookImportModalData { NgFor, AsyncPipe, NgOptimizedImage, + NgTemplateOutlet, ReactiveFormsModule, NzFormModule, NzSelectModule, @@ -56,6 +62,9 @@ export interface NotebookImportModalData { NzUploadComponent, NzButtonComponent, NzIconDirective, + NzTabsComponent, + NzTabComponent, + NzTabDirective, ], }) export class NotebookImportModalComponent implements OnDestroy { @@ -73,6 +82,25 @@ export class NotebookImportModalComponent implements OnDestroy { // and an empty list (no models available, e.g. the fetch failed or the feature is off). public readonly models$: Observable<{ name: string }[]> = this.notebookMigrationService.getAvailableModels(); + // The pane cross-fade reads as a flicker on a dense form, so only the ink bar animates. + // Held as a field rather than an inline literal so the binding keeps a stable reference. + public readonly tabAnimation = { inkBar: true, tabPane: false }; + + // Tab order: 0 = Jupyter notebook, 1 = Python file. + public selectedTabIndex = 0; + + /** + * Switching tabs drops the selected file: the two tabs accept different extensions, so + * carrying a selection across would leave, say, an .ipynb staged under the Python tab. + * The model stays selected because it applies to either input. + */ + public onTabChange(index: number): void { + this.selectedTabIndex = index; + const fileControl = this.importForm.get("file"); + fileControl?.reset(null); + fileControl?.updateValueAndValidity(); + } + public beforeUpload = (file: NzUploadFile) => { this.importForm.patchValue({ file }); this.importForm.get("file")?.markAsDirty(); diff --git a/frontend/src/app/workspace/service/notebook-migration/migration-llm.spec.ts b/frontend/src/app/workspace/service/notebook-migration/migration-llm.spec.ts index da1961efa3..e0e59fc9a4 100644 --- a/frontend/src/app/workspace/service/notebook-migration/migration-llm.spec.ts +++ b/frontend/src/app/workspace/service/notebook-migration/migration-llm.spec.ts @@ -208,6 +208,22 @@ describe("NotebookMigrationLLM", () => { warn.mockRestore(); }); + it("skips a mapping entry whose cell list is not an array, keeping the rest of the conversion", async () => { + const llm = makeLLM(); + mockResponses( + JSON.stringify({ code: { UDF1: "# UDF1", UDF2: "# UDF2" }, edges: [], outputs: {} }), + // A model that answers with a bare string here used to throw away the whole conversion. + JSON.stringify({ UDF1: "cell-a", UDF2: ["cell-b"] }) + ); + + const result = JSON.parse( + await llm.convertNotebookToWorkflow({ cells: [codeCell("cell-a", "x = 1"), codeCell("cell-b", "y = 2")] }) + ); + + expect(result.workflowJSON.operators).toHaveLength(2); + expect(result.workflowNotebookMapping.operator_to_cell).toEqual({ "PythonUDFV2-1": ["cell-b"] }); + }); + it("skips (with a warning) a mapping entry that references an unknown UDF id", async () => { const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); const notebook: Notebook = { cells: [codeCell("CELL1", "a")] }; @@ -522,4 +538,162 @@ describe("NotebookMigrationLLM", () => { expect(callModelSpy).not.toHaveBeenCalled(); }); }); + + describe("convertScriptToWorkflow", () => { + // Five lines; line 2 and line 4 are blank, so the reconciler's blank-gap handling shows up. + const script = ["import os", "", "x = compute()", "", "print(x)"].join("\n"); + + const workflowResponse = JSON.stringify({ + code: { UDF1: "# UDF1", UDF2: "# UDF2" }, + edges: [["UDF1", "UDF2"]], + outputs: { UDF1: ["value"], UDF2: ["result"] }, + }); + + // sendPrompt pushes onto the same array it hands to callModel, so a captured call argument + // keeps mutating. Each conversion does get a fresh array from seedDocumentation, so index by + // the conversion's first call and read the whole exchange rather than a point-in-time state. + function conversation(firstCallIndex = 0): { role: string; content: string }[] { + return callModelSpy.mock.calls[firstCallIndex][0] as { role: string; content: string }[]; + } + + function contentsFor(role: string, firstCallIndex = 0): string[] { + return conversation(firstCallIndex) + .filter(message => message.role === role) + .map(message => message.content); + } + + it("sends the script with a line number on every line", async () => { + const llm = makeLLM(); + mockResponses(workflowResponse, JSON.stringify({ UDF1: [[3, 3]] })); + + await llm.convertScriptToWorkflow(script); + + const prompt = contentsFor("user").join("\n"); + expect(prompt).toContain("1| import os"); + expect(prompt).toContain("3| x = compute()"); + expect(prompt).toContain("5| print(x)"); + }); + + it("seeds the script prelude, not the cell-marked notebook example", async () => { + const llm = makeLLM(); + mockResponses(workflowResponse, JSON.stringify({ UDF1: [[3, 3]] })); + + await llm.convertScriptToWorkflow(script); + + const prelude = contentsFor("system").join("\n"); + // The notebook worked example is the one entry that is swapped out; if it survives, the + // model is being shown cell ids for an input that has none. + expect(prelude).not.toContain("# START CELL1"); + expect(prelude).toContain("1| import pandas as pd"); + }); + + it("builds the workflow from the model's code and edges", async () => { + const llm = makeLLM(); + mockResponses(workflowResponse, JSON.stringify({ UDF1: [[3, 3]], UDF2: [[5, 5]] })); + + const { workflowJSON } = await llm.convertScriptToWorkflow(script); + + expect(workflowJSON.operators.map(operator => operator.operatorID)).toEqual(["PythonUDFV2-0", "PythonUDFV2-1"]); + expect(workflowJSON.links).toHaveLength(1); + expect(workflowJSON.links[0].source.operatorID).toBe("PythonUDFV2-0"); + expect(workflowJSON.links[0].target.operatorID).toBe("PythonUDFV2-1"); + }); + + it("derives cells from the reported ranges and maps them to operator ids", async () => { + const llm = makeLLM(); + mockResponses(workflowResponse, JSON.stringify({ UDF1: [[3, 3]], UDF2: [[5, 5]] })); + + const { workflowNotebookMapping, notebook } = await llm.convertScriptToWorkflow(script); + + // Line 4 is blank and is dropped; lines 1-2 are a real gap no UDF claimed. + expect(notebook.cells.map(cell => cell.source)).toEqual(["import os\n", "x = compute()", "print(x)"]); + + const [gap, first, second] = notebook.cells.map(cell => String(cell.metadata.uuid)); + expect(workflowNotebookMapping.operator_to_cell).toEqual({ + "PythonUDFV2-0": [first], + "PythonUDFV2-1": [second], + }); + expect(workflowNotebookMapping.cell_to_operator).toEqual({ + [first]: ["PythonUDFV2-0"], + [second]: ["PythonUDFV2-1"], + }); + // The unclaimed lines survive as a cell so no source is lost, but highlight nothing. + expect(workflowNotebookMapping.cell_to_operator[gap]).toBeUndefined(); + }); + + it("returns a notebook Jupyter can open", async () => { + const llm = makeLLM(); + mockResponses(workflowResponse, JSON.stringify({ UDF1: [[3, 3]] })); + + const { notebook } = await llm.convertScriptToWorkflow(script); + + expect(notebook.nbformat).toBe(4); + expect(notebook.nbformat_minor).toBe(4); + expect(notebook.metadata?.["language_info"]).toEqual({ name: "python" }); + notebook.cells.forEach(cell => { + expect(cell.cell_type).toBe("code"); + expect(String(cell.metadata.uuid)).not.toBe(""); + expect(cell.outputs).toEqual([]); + expect(cell.execution_count).toBeNull(); + }); + }); + + it("still returns the workflow when the model reports no usable ranges", async () => { + const llm = makeLLM(); + mockResponses(workflowResponse, JSON.stringify({ UDF1: "not a range" })); + + const { workflowJSON, workflowNotebookMapping, notebook } = await llm.convertScriptToWorkflow(script); + + expect(workflowJSON.operators).toHaveLength(2); + // Degrades to the whole script in one cell that highlights nothing, rather than failing + // and throwing away a conversion that already cost two model calls. + expect(notebook.cells).toHaveLength(1); + expect(notebook.cells[0].source).toBe(script); + expect(workflowNotebookMapping.operator_to_cell).toEqual({}); + }); + + it("skips a reported range whose UDF the model never defined", async () => { + const llm = makeLLM(); + mockResponses(workflowResponse, JSON.stringify({ UDF1: [[3, 3]], UDF9: [[5, 5]] })); + + const { workflowNotebookMapping } = await llm.convertScriptToWorkflow(script); + + expect(Object.keys(workflowNotebookMapping.operator_to_cell)).toEqual(["PythonUDFV2-0"]); + }); + + it("does not carry a previous notebook conversion's prelude into a script conversion", async () => { + const llm = makeLLM(); + mockResponses( + JSON.stringify({ code: { UDF1: "# UDF1" }, edges: [], outputs: {} }), + JSON.stringify({ UDF1: ["cell-a"] }) + ); + await llm.convertNotebookToWorkflow({ cells: [codeCell("cell-a", "x = 1")] }); + + mockResponses(workflowResponse, JSON.stringify({ UDF1: [[3, 3]] })); + await llm.convertScriptToWorkflow(script); + + // Calls 0 and 1 were the notebook conversion; the script conversion starts at call 2. + expect(contentsFor("system", 2).join("\n")).not.toContain("# START CELL1"); + // Nothing from the notebook exchange survives: not its cell ids, not its replies. + expect( + conversation(2) + .map(message => message.content) + .join("\n") + ).not.toContain("cell-a"); + }); + + it("refuses to run before initialize()", async () => { + const llm = makeUninitializedLLM(); + + await expect(llm.convertScriptToWorkflow(script)).rejects.toThrow("LLM session not initialized"); + expect(callModelSpy).not.toHaveBeenCalled(); + }); + + it("refuses to run when the feature is disabled", async () => { + const llm = makeUninitializedLLM(false); + + await expect(llm.convertScriptToWorkflow(script)).rejects.toThrow("Notebook migration feature is disabled"); + expect(callModelSpy).not.toHaveBeenCalled(); + }); + }); }); diff --git a/frontend/src/app/workspace/service/notebook-migration/migration-llm.ts b/frontend/src/app/workspace/service/notebook-migration/migration-llm.ts index 4bafa095e2..ff048a5f5a 100644 --- a/frontend/src/app/workspace/service/notebook-migration/migration-llm.ts +++ b/frontend/src/app/workspace/service/notebook-migration/migration-llm.ts @@ -38,20 +38,31 @@ import { EXAMPLE_OF_MULTIPLE_UDF_CONVERSION, WORKFLOW_PROMPT, MAPPING_PROMPT, + EXAMPLE_OF_MULTIPLE_UDF_CONVERSION_SCRIPT, + SCRIPT_WORKFLOW_PROMPT, + SCRIPT_MAPPING_PROMPT, } from "./migration-prompts"; +import { DerivedCell, segmentScript, splitScriptLines } from "./script-segmentation"; interface Cell { cell_type: string; metadata: { [key: string]: any }; // nbformat stores source as either a single string or an array of line strings. source: string | string[]; + outputs?: unknown[]; + execution_count?: number | null; } export interface Notebook { cells: Cell[]; + // Present in any real .ipynb and required by Jupyter's contents API, so the notebook + // synthesized for a script declares them too. + metadata?: { [key: string]: any }; + nbformat?: number; + nbformat_minor?: number; } -interface WorkflowJSON { +export interface WorkflowJSON { operators: OperatorPredicate[]; operatorPositions: Record<string, { x: number; y: number }>; links: any[]; @@ -59,20 +70,64 @@ interface WorkflowJSON { settings: WorkflowSettings; } -interface CombinedMapping { +export interface CombinedMapping { operator_to_cell: Record<string, string[]>; cell_to_operator: Record<string, string[]>; } /** - * Wraps a single LLM chat session that converts a Jupyter notebook into a Texera - * workflow plus a cell<->operator mapping. + * A script conversion also yields the notebook it derived, because nothing upstream had one: + * the caller stores and displays it exactly as it would a user's own .ipynb. + */ +export interface ScriptConversion { + workflowJSON: WorkflowJSON; + workflowNotebookMapping: CombinedMapping; + notebook: Notebook; +} + +// Prefix each line with its number so the model can report ranges without counting lines itself. +// Uses the segmenter's own line splitting: the reported numbers only mean anything if the side +// that numbers and the side that slices agree on what a line is. +function numberScriptLines(source: string): string { + const lines = splitScriptLines(source); + const width = String(lines.length).length; + return lines.map((line, index) => `${String(index + 1).padStart(width, " ")}| ${line}`).join("\n"); +} + +// Wrap derived cells as a notebook. The nbformat fields are what make it openable in Jupyter, +// and metadata.uuid is the join key the stored mapping is expressed in, same as for a real .ipynb. +function toDerivedNotebook(cells: DerivedCell[]): Notebook { + return { + cells: cells.map(cell => ({ + cell_type: "code", + metadata: { uuid: cell.uuid }, + source: cell.source, + outputs: [], + execution_count: null, + })), + metadata: { + kernelspec: { display_name: "Python 3", language: "python", name: "python3" }, + language_info: { name: "python" }, + }, + nbformat: 4, + nbformat_minor: 4, + }; +} + +/** + * Wraps a single LLM chat session that converts a Jupyter notebook or a Python script into + * a Texera workflow plus a cell<->operator mapping. * * Lifecycle: `initialize()` -> `verifyConnection()` (optional) -> - * `convertNotebookToWorkflow()` -> `close()`. The session keeps a running `messages` - * history shared by the prompts within one conversion. `convertNotebookToWorkflow()` - * resets that history to the documentation prelude at its start, so the same instance - * can convert multiple notebooks without leaking one conversion's context into the next. + * `convertNotebookToWorkflow()` or `convertScriptToWorkflow()` -> `close()`. The session keeps + * a running `messages` history shared by the prompts within one conversion. Each conversion + * resets that history to its documentation prelude at its start, so the same instance can run + * several conversions, in either mode, without leaking one conversion's context into the next. + * + * The two modes differ only in framing. A notebook arrives already split into cells and the + * model is asked to map UDFs onto those cell ids; a script has no cells, so it is sent with + * line numbers, the model reports the line ranges each UDF came from, and the cells are derived + * from that answer. Everything downstream of the model's reply is shared. * * Output column types: intermediate UDFs declare their output columns as `binary` so rich * Python objects (DataFrames, arrays, models) round-trip between operators via pickle. @@ -109,6 +164,14 @@ export class NotebookMigrationLLM { EXAMPLE_OF_MULTIPLE_UDF_CONVERSION, ]; + // The script prelude differs in exactly one entry. The notebook worked example is written in + // `# START CELL1` form, and a system-message example of that weight would push the model to + // answer in cell ids for an input that has no cells. The notebook array is left untouched so + // existing conversions see byte-identical context. + private static readonly SCRIPT_DOCUMENTATION: string[] = NotebookMigrationLLM.DOCUMENTATION.map(doc => + doc === EXAMPLE_OF_MULTIPLE_UDF_CONVERSION ? EXAMPLE_OF_MULTIPLE_UDF_CONVERSION_SCRIPT : doc + ); + constructor( private config: GuiConfigService, private workflowUtilService: WorkflowUtilService @@ -125,11 +188,12 @@ export class NotebookMigrationLLM { } /** - * Seed the conversation with the Texera documentation prelude, discarding any - * prior conversation. Used by initialize() and at the start of each conversion. + * Seed the conversation with a Texera documentation prelude, discarding any prior + * conversation. Used by initialize() and at the start of each conversion, which is where + * the input-specific variant is chosen. */ - private seedDocumentation(): void { - this.messages = NotebookMigrationLLM.DOCUMENTATION.map( + private seedDocumentation(documentation: string[] = NotebookMigrationLLM.DOCUMENTATION): void { + this.messages = documentation.map( (doc): ModelMessage => ({ role: "system", content: doc, @@ -269,7 +333,7 @@ export class NotebookMigrationLLM { // Reset to the documentation prelude so a prior conversion's prompts/responses // don't leak into this one. The two sendPrompt calls below still share history. - this.seedDocumentation(); + this.seedDocumentation(NotebookMigrationLLM.DOCUMENTATION); const codeCells = notebook.cells.filter(cell => cell.cell_type === "code"); @@ -294,7 +358,61 @@ export class NotebookMigrationLLM { // Remove ```json blocks and parse const udfLLMResponse = this.parseJsonResponse(workflow, "workflow"); + const { workflowJSON, udfIdToOperatorId } = this.buildWorkflow(udfLLMResponse); + // The notebook path keys its mapping on the cell uuids embedded in the prompt. + const parsedMapping: Record<string, string[]> = this.parseJsonResponse(mapping, "mapping"); + const workflowNotebookMapping = this.buildCombinedMapping(parsedMapping, udfIdToOperatorId); + + return JSON.stringify({ workflowJSON, workflowNotebookMapping }); + } + + /** + * Send a Python script to be converted into a workflow, a mapping, and the notebook the + * mapping is expressed against. + * + * Differs from the notebook path in two places only. The script is sent with line numbers + * rather than cell markers, and the model is asked which line ranges became which UDF; the + * cells are then derived from that answer instead of arriving with the input. + */ + public async convertScriptToWorkflow(source: string): Promise<ScriptConversion> { + this.assertEnabled(); + if (!this.initialized) { + throw new Error("LLM session not initialized"); + } + + this.seedDocumentation(NotebookMigrationLLM.SCRIPT_DOCUMENTATION); + + const workflow = await this.sendPrompt(`${SCRIPT_WORKFLOW_PROMPT}\n${numberScriptLines(source)}`); + const mapping = await this.sendPrompt(SCRIPT_MAPPING_PROMPT); + + const udfLLMResponse = this.parseJsonResponse(workflow, "workflow"); + const { workflowJSON, udfIdToOperatorId } = this.buildWorkflow(udfLLMResponse); + + // segmentScript reconciles whatever the model reported, so a malformed range degrades the + // mapping rather than discarding a workflow that already cost a full conversion. + const reportedRanges = this.parseJsonResponse(mapping, "mapping"); + const { cells, udfToCellUuids } = segmentScript(source, reportedRanges); + + return { + workflowJSON, + workflowNotebookMapping: this.buildCombinedMapping(udfToCellUuids, udfIdToOperatorId), + notebook: toDerivedNotebook(cells), + }; + } + + /** + * Assemble the workflow from the model's `code`, `edges` and `outputs` response. + * + * Input-agnostic: the notebook and script paths differ in how they prompt and in what their + * mapping is keyed on, not in how the generated UDFs become operators. + * + * Returns the workflow together with the UDF id -> operatorID index the mapping is built from. + */ + private buildWorkflow(udfLLMResponse: any): { + workflowJSON: WorkflowJSON; + udfIdToOperatorId: Record<string, string>; + } { const workflowJSON: WorkflowJSON = { operators: [], operatorPositions: {}, @@ -306,7 +424,7 @@ export class NotebookMigrationLLM { }, }; - const udfMappingToUUID: Record<string, string> = {}; + const udfIdToOperatorId: Record<string, string> = {}; // UDFs that are never the source of an edge are terminal (result-facing). Their outputs // default to "string" so the result panel renders typed values; intermediate UDFs keep @@ -336,12 +454,12 @@ export class NotebookMigrationLLM { }, }; - udfMappingToUUID[udfId] = operator.operatorID; + udfIdToOperatorId[udfId] = operator.operatorID; workflowJSON.operators.push(operator); workflowJSON.operatorPositions[operator.operatorID] = { x: 140 * (i + 1), y: 0 }; }); - const knownUdfIds = new Set(Object.keys(udfMappingToUUID)); + const knownUdfIds = new Set(Object.keys(udfIdToOperatorId)); // Add links/edges. Skip (with a warning) any edge that references a UDF id the LLM // never defined in `code`, rather than emitting a link with an undefined endpoint. @@ -353,44 +471,55 @@ export class NotebookMigrationLLM { workflowJSON.links.push({ linkID: `link-${uuidv4()}`, source: { - operatorID: udfMappingToUUID[source], + operatorID: udfIdToOperatorId[source], portID: "output-0", }, target: { - operatorID: udfMappingToUUID[target], + operatorID: udfIdToOperatorId[target], portID: "input-0", }, }); }); - // Parse mapping - const parsedMapping: Record<string, string[]> = this.parseJsonResponse(mapping, "mapping"); - - const udfToCell: Record<string, string[]> = {}; - const cellToUdf: Record<string, string[]> = {}; + return { workflowJSON, udfIdToOperatorId }; + } - Object.entries(parsedMapping).forEach(([udf, cells]) => { - if (!knownUdfIds.has(udf)) { - console.warn(`Skipping mapping entry with unknown UDF id: ${udf}`); + /** + * Invert a UDF id -> cell ids mapping into the stored operator<->cell form, skipping (with a + * warning) any UDF the model never defined in `code`. Shared by both input paths: they differ + * only in where the cell ids came from. + */ + private buildCombinedMapping( + udfToCells: Record<string, string[]>, + udfIdToOperatorId: Record<string, string> + ): CombinedMapping { + const operatorToCell: Record<string, string[]> = {}; + const cellToOperator: Record<string, string[]> = {}; + + Object.entries(udfToCells).forEach(([udfId, cells]) => { + const operatorId = udfIdToOperatorId[udfId]; + if (!operatorId) { + console.warn(`Skipping mapping entry with unknown UDF id: ${udfId}`); return; } - const udfUUID = udfMappingToUUID[udf]; - udfToCell[udfUUID] = cells; + // The notebook path's cell ids come straight from an unvalidated model reply. Without this + // a non-array would throw from forEach below and discard a conversion that already cost two + // model calls; skipping degrades the mapping instead, as the script path already does. + if (!Array.isArray(cells)) { + console.warn(`Skipping mapping entry whose cell list is not an array, for UDF id: ${udfId}`); + return; + } + operatorToCell[operatorId] = cells; cells.forEach(cell => { - if (!cellToUdf[cell]) { - cellToUdf[cell] = [udfUUID]; + if (!cellToOperator[cell]) { + cellToOperator[cell] = [operatorId]; } else { - cellToUdf[cell].push(udfUUID); + cellToOperator[cell].push(operatorId); } }); }); - const workflowNotebookMapping: CombinedMapping = { - operator_to_cell: udfToCell, - cell_to_operator: cellToUdf, - }; - - return JSON.stringify({ workflowJSON, workflowNotebookMapping }); + return { operator_to_cell: operatorToCell, cell_to_operator: cellToOperator }; } /** diff --git a/frontend/src/app/workspace/service/notebook-migration/migration-prompts.ts b/frontend/src/app/workspace/service/notebook-migration/migration-prompts.ts index 2594d2d58e..43f0d9dea1 100644 --- a/frontend/src/app/workspace/service/notebook-migration/migration-prompts.ts +++ b/frontend/src/app/workspace/service/notebook-migration/migration-prompts.ts @@ -364,7 +364,7 @@ It is VERY important that all of the original code in the Jupyter notebook is re Make sure that nothing in the original is removed and that the semantic meaning of what the original code was doing is retained. The only exception is data-loading code (e.g. pd.read_csv); it is represented by the workflow's input/source operator rather than copied into a UDF. If there are user-defined Python classes, include the entire class definition in the appropriate UDF(s) that use that class. -Always include the code that defines the class inside of every distinct UDF that uses that constructs an object of that class. +Always include the full class definition inside every UDF that references that class, including every UDF that constructs an object of it. Python classes are allowed in Texera UDFs and follow the same semantics as standard Python. They can be defined outside of ProcessTableOperator, ProcessTupleOperator, and ProcessBatchOperator. @@ -408,7 +408,181 @@ Here is an example of a mapping generated between the given example Python code ], "UDF4": [ "CELL8" +], +"UDF5": [ +"CELL9" ] } Now create a mapping for the UDFs and the original code. Link the code blocks marked by 'START <cell-uuid>' and 'END <cell-uuid>' with the UDF UUID's. The code between them should be equivalent. Multiple cells can be mapped to the same UDF when that UDF implements the logic of those cells. There could be any number of cells and UDFs, so only create the correct number in the mapping. Only give the mapping. `; + +export const EXAMPLE_OF_MULTIPLE_UDF_CONVERSION_SCRIPT = ` +Here is an example of breaking up python code into multiple Texera UDFs. Format your response structure exactly like the given example. The "code" key contains a dictionary of the UDF ID's with their respective code. The "edges" key contains a list of pairs that contains the connections between UDFs. The "outputs" key contains a dictionary of the UDF ID's with a list of the output column names of the DataFrame that the UDF yields. The UDFs can branch and merge, it does not have to be a l [...] + +The original code is shown with each line prefixed by its line number and a '|'. Those prefixes are annotations so that line ranges can be referred to later. They are not part of the code and must never appear in the code you generate. + +Original Code: +\`\`\`python + 1| import pandas as pd + 2| from sklearn.model_selection import train_test_split + 3| from sklearn.ensemble import RandomForestClassifier + 4| from sklearn.svm import SVC + 5| from sklearn.tree import DecisionTreeClassifier + 6| from sklearn.linear_model import LogisticRegression + 7| from sklearn.metrics import accuracy_score + 8| from sklearn.preprocessing import StandardScaler + 9| import matplotlib.pyplot as plt +10| +11| # Load the dataset +12| file_path = 'diabetes.csv' +13| data = pd.read_csv(file_path) +14| +15| # Remove duplicate rows +16| data = data.drop_duplicates() +17| +18| # Remove rows with null values +19| data = data.dropna() +20| +21| # Print the minimum, maximum, and mean for all fields +22| print("Minimum values:", data.min()) +23| print("Maximum values:", data.max()) +24| print("Mean values:", data.mean()) +25| +26| # Create a boxplot for the 'Pregnancies' field +27| plt.figure(figsize=(8, 6)) +28| plt.boxplot(data['Pregnancies'], vert=False, patch_artist=True) +29| plt.title('Boxplot of Pregnancies') +30| plt.xlabel('Number of Pregnancies') +31| plt.show() +32| +33| # Separate features and target variable +34| X = data.drop('Outcome', axis=1) +35| y = data['Outcome'] +36| +37| # Split data into training and testing sets (80% train, 20% test) +38| X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) +39| +40| scaler = StandardScaler() +41| X_train = scaler.fit_transform(X_train) +42| X_test = scaler.transform(X_test) +43| +44| # Train Random Forest model +45| rf_model = RandomForestClassifier(random_state=42) +46| rf_model.fit(X_train, y_train) +47| rf_pred = rf_model.predict(X_test) +48| rf_accuracy = accuracy_score(y_test, rf_pred) +49| print(f"Random Forest Accuracy: {rf_accuracy:.2%}") +50| +51| # Train SVM model +52| svm_model = SVC(random_state=42) +53| svm_model.fit(X_train, y_train) +54| svm_pred = svm_model.predict(X_test) +55| svm_accuracy = accuracy_score(y_test, svm_pred) +56| print(f"SVM Accuracy: {svm_accuracy:.2%}") +\`\`\` + +Texera UDF conversion: +\`\`\`json +{ + "code": { + "UDF1": "# UDF1\nfrom pytexera import *\nimport pandas as pd\nfrom typing import Iterator, Optional\n\nclass ProcessTableOperator(UDFTableOperator):\n\n @overrides\n def process_table(self, table: Table, port: int) -> Iterator[Optional[TableLike]]:\n # Remove duplicate rows\n data = table.drop_duplicates()\n\n # Remove rows with null values\n data = data.dropna()\n\n # Calculate statistics\n min_values = data.min()\n max_valu [...] + "UDF2": "# UDF2\nfrom pytexera import *\nimport pandas as pd\nimport plotly.express as px\nimport plotly.io\nfrom typing import Iterator, Optional\n\nclass ProcessTableOperator(UDFTableOperator):\n def render_error(self, error_msg):\n return '''<h1>Boxplot is not available.</h1>\n <p>Reason is: {} </p>\n '''.format(error_msg)\n\n @overrides\n def process_table(self, table: Table, port: int) -> Iterator[Optional[TableLike]]:\n [...] + "UDF3": "# UDF3\nfrom pytexera import *\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom typing import Iterator, Optional\n\nclass ProcessTableOperator(UDFTableOperator):\n\n @overrides\n def process_table(self, table: Table, port: int) -> Iterator[Optional[TableLike]]:\n data = table['data'].iloc[0]\n\n # Separate features and target variable\n X = data.drop('Outcome', ax [...] + "UDF4": "# UDF4\nfrom pytexera import *\nimport pandas as pd\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import accuracy_score\nfrom typing import Iterator, Optional\n\nclass ProcessTableOperator(UDFTableOperator):\n\n @overrides\n def process_table(self, table: Table, port: int) -> Iterator[Optional[TableLike]]:\n X_train = table['X_train'].iloc[0]\n y_train = table['y_train'].iloc[0]\n X_test = table['X_test'].iloc[0]\n [...] + "UDF5": "# UDF5\nfrom pytexera import *\nimport pandas as pd\nfrom sklearn.svm import SVC\nfrom sklearn.metrics import accuracy_score\nfrom typing import Iterator, Optional\n\nclass ProcessTableOperator(UDFTableOperator):\n\n @overrides\n def process_table(self, table: Table, port: int) -> Iterator[Optional[TableLike]]:\n X_train = table['X_train'].iloc[0]\n y_train = table['y_train'].iloc[0]\n X_test = table['X_test'].iloc[0]\n y_test = table['y [...] + }, + "edges": [ + ["UDF1", "UDF2"], + ["UDF1", "UDF3"], + ["UDF3", "UDF4"], + ["UDF3", "UDF5"] + ], + "outputs": { + "UDF1": ["min_values", "max_values", "mean_values", "data"], + "UDF2": ["html-content"], + "UDF3": ["X_train", "X_test", "y_train", "y_test"], + "UDF4": ["rf_model", "rf_accuracy", "X_test", "y_test"], + "UDF5": ["svm_model", "svm_accuracy", "X_test", "y_test"] + } +} +\`\`\``; + +export const SCRIPT_WORKFLOW_PROMPT = `You are an expert in Python coding and workflow systems. +Many users of Texera system are non-technical, but the Python scripts they provide are written by technical people. +They want to convert their scripts to Texera workflows. +Your goal is to help convert these scripts into a Texera workflow that non-technical users can use directly. +So do not remove or modify any classes or functions, preserve their names and structure as they are. +Ensure that all essential logic remains intact. +Create multiple Texera UDF codes using the provided Python code. +Number each UDF, starting at 1 and incrementing, by starting with a comment that states that UDF number. + +Use the class and function names as shown in ProcessTupleOperator, ProcessTableOperator, and ProcessBatchOperator. +Do not change the class names, function names, or input parameters. +Use the ones that make sense and split the code meaningfully as instructed. + +Use the starter code provided for Python UDFs. + +Use the documentation of Table, Tuple, or Batch to work with parameters within Texera UDF. +Do not import other libraries to define these types. + +There is no need for an __init__ function. Assume all inputs are valid pandas DataFrames, +so do not use .to_pandas(), .to_dataframe(), etc. Do not load data from a file in the first UDF; +the workflow's source operator supplies the initial data, so assume it is already given to you in the +table parameter. Replacing file-loading code with this input is the one exception to preserving all +original code (see below). +Ensure proper data flow between functions. Separate operators as if they will run in different files. + +Current UDF operators can only have one output. Build a dataframe to yield all necessary variables +and data. Ensure proper data flow for each UDF and all information is yielded (including training +and testing data) if subsequent UDFs need them. + +Ensure all necessary imports are included in each UDF code block. + +Each UDF operator should be in its own Python code block. Do not combine them into a single block. +Ensure import statements cover all used functions and separate them as necessary. + +It is VERY important that all of the original code in the Python script is represented in the generated workflow. +Make sure that nothing in the original is removed and that the semantic meaning of what the original code was doing is retained. +The only exception is data-loading code (e.g. pd.read_csv); it is represented by the workflow's input/source operator rather than copied into a UDF. +If there are user-defined Python classes, include the entire class definition in the appropriate UDF(s) that use that class. +Always include the full class definition inside every UDF that references that class, including every UDF that constructs an object of it. +Python classes are allowed in Texera UDFs and follow the same semantics as standard Python. +They can be defined outside of ProcessTableOperator, ProcessTupleOperator, and ProcessBatchOperator. + +Return only the JSON formatted response, do not give any explanation. +Do not wrap the JSON in markdown code fences. Output raw JSON only. +Make sure the response is a valid JSON structure, including closing all braces and not including commas after the last element. +Follow this JSON format (don't reuse the values, this is just the format). 'code', 'edges', and 'outputs' are all their own key's, do not nest any of these in another one and make sure to close their braces: +{ +"code": { +"UDF1": "code for UDF1 goes here", +"UDF2": "code for UDF2 goes here" +}, +"edges": [ +["UDF1", "UDF2"] +], +"outputs": { +"UDF1": ["min_values", "max_values", "mean_values", "data"], +"UDF2": ["html-content"] +} +} +Make sure only the keys in the code section appear in the edges and outputs sections. Do not include any extraneous fields. +Do not include any extraneous UDF's in the code field that include empty strings. +Give ALL of the code, do not omit anything or use placeholders for code. Make sure ALL code in the original is translated over. +The value of each UDF must be a valid JSON string: escape newlines, quotes, and backslashes correctly so that the decoded string is runnable Python. Use whichever quotes the Python code requires. +Each line of the script below is prefixed with its line number followed by '| '. Those prefixes are annotations +so that line ranges can be referred to later; never reproduce them in any generated UDF code. +Convert following the instructions and examples given. Here is the code: +`; + +export const SCRIPT_MAPPING_PROMPT = ` +Here is an example of a mapping generated between the given example Python code and the Texera UDFs, using line ranges of the original script and the UDF IDs. A range is a pair [firstLine, lastLine]; both bounds are 1-indexed and inclusive, and they refer to the line numbers shown in the prefix of the original code. A UDF may list several ranges when its logic came from separate parts of the script. The format should be kept the same. +{ +"UDF1": [[15, 24]], +"UDF2": [[26, 31]], +"UDF3": [[33, 42]], +"UDF4": [[44, 49]], +"UDF5": [[51, 56]] +} +Now create a mapping for the UDFs and the original code you were given. For each UDF, report the line ranges of the original script whose logic that UDF implements. The code in those lines should be equivalent to what the UDF does. Lines that no UDF implements, such as imports or the data loading that the workflow's source operator replaces, can be left out entirely. Give the first line before the last within each range, and do not shift the numbers: they must match the prefixes you were [...] +`; diff --git a/frontend/src/app/workspace/service/notebook-migration/notebook-migration.service.spec.ts b/frontend/src/app/workspace/service/notebook-migration/notebook-migration.service.spec.ts index 3dbe846203..a58856525d 100644 --- a/frontend/src/app/workspace/service/notebook-migration/notebook-migration.service.spec.ts +++ b/frontend/src/app/workspace/service/notebook-migration/notebook-migration.service.spec.ts @@ -371,6 +371,63 @@ describe("NotebookMigrationService", () => { }); }); + // sendScriptToAIGenerateWorkflow — same lifecycle as the notebook path, but the conversion + // hands back a derived notebook too, since a script arrives without one. + describe("sendScriptToAIGenerateWorkflow (enabled)", () => { + let fakeLLM: { + initialize: ReturnType<typeof vi.fn>; + verifyConnection: ReturnType<typeof vi.fn>; + convertScriptToWorkflow: ReturnType<typeof vi.fn>; + close: ReturnType<typeof vi.fn>; + }; + + const conversion = { + workflowJSON: { ops: 1 }, + workflowNotebookMapping: { m: 2 }, + notebook: { cells: [{ cell_type: "code", metadata: { uuid: "u1" }, source: "x = 1" }] }, + }; + + beforeEach(() => { + fakeLLM = { + initialize: vi.fn(), + verifyConnection: vi.fn().mockResolvedValue(true), + convertScriptToWorkflow: vi.fn(), + close: vi.fn(), + }; + vi.spyOn(service as any, "createMigrationLLM").mockReturnValue(fakeLLM); + }); + + it("returns the workflow, mapping and derived notebook, and closes the client", async () => { + fakeLLM.convertScriptToWorkflow.mockResolvedValue(conversion); + + const result = await service.sendScriptToAIGenerateWorkflow("x = 1", "gpt-4"); + + expect(result).toEqual({ + workflowContent: { ops: 1 }, + mappingContent: { m: 2 }, + notebook: conversion.notebook, + }); + expect(fakeLLM.convertScriptToWorkflow).toHaveBeenCalledWith("x = 1"); + expect(fakeLLM.initialize).toHaveBeenCalledWith("gpt-4"); + expect(fakeLLM.close).toHaveBeenCalled(); + }); + + it("rejects when the connection cannot be verified, and still closes the client", async () => { + fakeLLM.verifyConnection.mockResolvedValue(false); + + await expect(service.sendScriptToAIGenerateWorkflow("x = 1", "gpt-4")).rejects.toThrow(/authenticate/i); + expect(fakeLLM.convertScriptToWorkflow).not.toHaveBeenCalled(); + expect(fakeLLM.close).toHaveBeenCalled(); + }); + + it("rethrows conversion errors and still closes the client", async () => { + fakeLLM.convertScriptToWorkflow.mockRejectedValue(new Error("conversion boom")); + + await expect(service.sendScriptToAIGenerateWorkflow("x = 1", "gpt-4")).rejects.toThrow(/conversion boom/); + expect(fakeLLM.close).toHaveBeenCalled(); + }); + }); + // Feature flag gate (defence in depth). With the flag off, every public // method must short-circuit — no HTTP traffic, no fetch, no LLM lifecycle, // no notifications. @@ -389,6 +446,10 @@ describe("NotebookMigrationService", () => { await expect(service.sendToAIGenerateWorkflow({ cells: [] } as any, "gpt-4")).rejects.toThrow(/disabled/i); }); + it("sendScriptToAIGenerateWorkflow rejects with a disabled-feature error", async () => { + await expect(service.sendScriptToAIGenerateWorkflow("x = 1", "gpt-4")).rejects.toThrow(/disabled/i); + }); + it("sendNotebookToJupyter returns 0 with no HTTP call or notification", async () => { const result = await service.sendNotebookToJupyter({ cells: [] } as any, "notebook_1.ipynb"); expect(result).toBe(0); @@ -472,4 +533,37 @@ describe("NotebookMigrationService", () => { await expect(service.parseAndTagNotebook(new File([""], "x.ipynb"))).rejects.toThrow(/Failed to read/); }); }); + + // parseScriptFile (reads an uploaded .py as text) + describe("parseScriptFile", () => { + it("resolves the file's contents verbatim", async () => { + const source = "import os\n\nprint(os.getcwd())\n"; + + await expect(service.parseScriptFile(new File([source], "script.py"))).resolves.toBe(source); + }); + + it("rejects an empty file before it can cost an LLM round trip", async () => { + await expect(service.parseScriptFile(new File([""], "empty.py"))).rejects.toThrow(/empty/i); + }); + + it("rejects a file holding only whitespace", async () => { + await expect(service.parseScriptFile(new File([" \n\t\n"], "blank.py"))).rejects.toThrow(/empty/i); + }); + + it("rejects when the file content is not a string", async () => { + vi.spyOn(FileReader.prototype, "readAsText").mockImplementation(function (this: any) { + // result is a getter-only property, so shadow it with an own non-string value. + Object.defineProperty(this, "result", { value: null, configurable: true }); + this.onload?.(new ProgressEvent("load")); + }); + await expect(service.parseScriptFile(new File([""], "x.py"))).rejects.toThrow(/not a valid string/); + }); + + it("rejects when the file cannot be read", async () => { + vi.spyOn(FileReader.prototype, "readAsText").mockImplementation(function (this: any) { + this.onerror?.(new ProgressEvent("error")); + }); + await expect(service.parseScriptFile(new File([""], "x.py"))).rejects.toThrow(/Failed to read/); + }); + }); }); diff --git a/frontend/src/app/workspace/service/notebook-migration/notebook-migration.service.ts b/frontend/src/app/workspace/service/notebook-migration/notebook-migration.service.ts index ed3ffd5a6e..61244f1232 100644 --- a/frontend/src/app/workspace/service/notebook-migration/notebook-migration.service.ts +++ b/frontend/src/app/workspace/service/notebook-migration/notebook-migration.service.ts @@ -103,11 +103,59 @@ export class NotebookMigrationService { notebookContent: Notebook, modelType: string ): Promise<{ workflowContent: WorkflowContent; mappingContent: MappingContent }> { + return this.withMigrationLLM(modelType, async migrationLLM => { + try { + const result = await migrationLLM.convertNotebookToWorkflow(notebookContent); + const parsedResult = JSON.parse(result); + const workflowContent = parsedResult.workflowJSON; + const mappingContent = parsedResult.workflowNotebookMapping; + return { workflowContent, mappingContent }; + } catch (error) { + console.error("Error converting notebook:", error); + throw error; + } + }); + } + + /** + * Convert a Python script into a workflow. + * + * Returns a notebook alongside the workflow and mapping, which the notebook path does not: + * a script has no cells, so the LLM reports which line ranges became which operator and the + * notebook is derived from that. Callers store and display it as they would a user's own + * .ipynb, so the Jupyter panel and cell highlighting work the same way for both inputs. + */ + public async sendScriptToAIGenerateWorkflow( + scriptSource: string, + modelType: string + ): Promise<{ workflowContent: WorkflowContent; mappingContent: MappingContent; notebook: Notebook }> { + return this.withMigrationLLM(modelType, async migrationLLM => { + try { + const conversion = await migrationLLM.convertScriptToWorkflow(scriptSource); + return { + workflowContent: conversion.workflowJSON, + mappingContent: conversion.workflowNotebookMapping, + notebook: conversion.notebook, + }; + } catch (error) { + console.error("Error converting Python script:", error); + throw error; + } + }); + } + + /** + * Run one conversion against a fresh LLM session: initialize, check the backend is reachable, + * then convert. Shared by both input paths so the lifecycle cannot drift between them, and so + * the finally guarantees close() for every exit, including a failed verifyConnection. + */ + private async withMigrationLLM<T>( + modelType: string, + convert: (migrationLLM: NotebookMigrationLLM) => Promise<T> + ): Promise<T> { if (!this.enabled) throw new Error("Notebook migration feature is disabled"); const migrationLLM = this.createMigrationLLM(); // initialize() defaults to the user's Texera JWT via AuthService.getAccessToken(). - // The outer try/finally guarantees close() runs for the whole lifecycle, - // including a verifyConnection failure. try { migrationLLM.initialize(modelType); @@ -116,16 +164,7 @@ export class NotebookMigrationService { throw new Error("Unable to authenticate with or reach the LLM backend"); } - try { - const result = await migrationLLM.convertNotebookToWorkflow(notebookContent); - const parsedResult = JSON.parse(result); - const workflowContent = parsedResult.workflowJSON; - const mappingContent = parsedResult.workflowNotebookMapping; - return { workflowContent, mappingContent }; - } catch (error) { - console.error("Error converting notebook:", error); - throw error; - } + return await convert(migrationLLM); } finally { migrationLLM.close(); } @@ -279,6 +318,28 @@ export class NotebookMigrationService { delete this.mapping[id]; } + // Reads a .py file as text. Rejects on a read error or a file with nothing in it: an empty + // script would otherwise cost a full LLM round trip to produce an empty workflow. Uses + // FileReader for the same reason parseAndTagNotebook does. + public parseScriptFile(file: File): Promise<string> { + return new Promise((resolve, reject) => { + const reader = new FileReader(); + reader.onerror = () => reject(new Error("Failed to read the Python file.")); + reader.onload = () => { + if (typeof reader.result !== "string") { + reject(new Error("File content is not a valid string.")); + return; + } + if (reader.result.trim() === "") { + reject(new Error("The Python file is empty.")); + return; + } + resolve(reader.result); + }; + reader.readAsText(file); + }); + } + // Reads and parses an .ipynb file, then tags each cell with a uuid (the mapping keys off these). // Rejects on a read error, invalid JSON, or a missing cells array. Uses FileReader rather than // file.text() because jsdom (the test environment) does not implement Blob/File.text(). diff --git a/frontend/src/app/workspace/service/notebook-migration/script-segmentation.spec.ts b/frontend/src/app/workspace/service/notebook-migration/script-segmentation.spec.ts new file mode 100644 index 0000000000..ae1df858b4 --- /dev/null +++ b/frontend/src/app/workspace/service/notebook-migration/script-segmentation.spec.ts @@ -0,0 +1,350 @@ +/** + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, + * software distributed under the License is distributed on an + * "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY + * KIND, either express or implied. See the License for the + * specific language governing permissions and limitations + * under the License. + */ + +import { segmentScript } from "./script-segmentation"; + +describe("segmentScript", () => { + // Sequential ids keep assertions readable; the real factory is uuidv4. + function counterIds(): () => string { + let next = 0; + return () => `c${++next}`; + } + + function segment(source: string, ranges: Record<string, unknown> | null | undefined) { + return segmentScript(source, ranges, counterIds()); + } + + // A 10-line script; every line is distinguishable so slices can be asserted exactly. + const script = ["L1", "L2", "L3", "L4", "L5", "L6", "L7", "L8", "L9", "L10"].join("\n"); + + // Cells reduced to the shape the assertions care about. + function spans(result: ReturnType<typeof segment>): [number, number][] { + return result.cells.map(cell => [cell.startLine, cell.endLine]); + } + + describe("well-formed input", () => { + it("maps a single range covering the whole file to one cell", () => { + const result = segment(script, { UDF1: [[1, 10]] }); + + expect(spans(result)).toEqual([[1, 10]]); + expect(result.cells[0].source).toBe(script); + expect(result.udfToCellUuids).toEqual({ UDF1: ["c1"] }); + }); + + it("splits adjacent ranges into one cell each, in source order", () => { + const result = segment(script, { UDF1: [[1, 4]], UDF2: [[5, 10]] }); + + expect(spans(result)).toEqual([ + [1, 4], + [5, 10], + ]); + expect(result.cells[0].source).toBe("L1\nL2\nL3\nL4"); + expect(result.cells[1].source).toBe("L5\nL6\nL7\nL8\nL9\nL10"); + expect(result.udfToCellUuids).toEqual({ UDF1: ["c1"], UDF2: ["c2"] }); + }); + + it("gives a UDF every cell its several ranges cover", () => { + const result = segment(script, { UDF1: [[1, 3]], UDF2: [[4, 6]] }); + const both = segment(script, { + UDF1: [ + [1, 3], + [7, 10], + ], + UDF2: [[4, 6]], + }); + + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1"]); + expect(both.udfToCellUuids["UDF1"]).toEqual(["c1", "c3"]); + expect(both.udfToCellUuids["UDF2"]).toEqual(["c2"]); + }); + }); + + describe("gaps and overlaps", () => { + it("keeps a gap between two ranges as its own cell that no UDF claims", () => { + const result = segment(script, { UDF1: [[1, 3]], UDF2: [[8, 10]] }); + + expect(spans(result)).toEqual([ + [1, 3], + [4, 7], + [8, 10], + ]); + // The gap's source survives even though nothing maps to it, so no code is lost. + expect(result.cells[1].source).toBe("L4\nL5\nL6\nL7"); + expect(Object.values(result.udfToCellUuids).flat()).not.toContain("c2"); + }); + + it("splits an overlap into a shared cell that maps to both UDFs", () => { + const result = segment(script, { + UDF1: [[1, 6]], + UDF2: [[4, 10]], + }); + + expect(spans(result)).toEqual([ + [1, 3], + [4, 6], + [7, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1", "c2"]); + expect(result.udfToCellUuids["UDF2"]).toEqual(["c2", "c3"]); + }); + + it("handles a range fully nested inside another", () => { + const result = segment(script, { UDF1: [[1, 10]], UDF2: [[4, 6]] }); + + expect(spans(result)).toEqual([ + [1, 3], + [4, 6], + [7, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1", "c2", "c3"]); + expect(result.udfToCellUuids["UDF2"]).toEqual(["c2"]); + }); + + it("emits cells in source order even when the ranges arrive out of order", () => { + const result = segment(script, { UDF1: [[7, 10]], UDF2: [[1, 3]] }); + + expect(spans(result)).toEqual([ + [1, 3], + [4, 6], + [7, 10], + ]); + expect(result.udfToCellUuids["UDF2"]).toEqual(["c1"]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c3"]); + }); + + it("collapses duplicate ranges for the same UDF into one cell", () => { + const result = segment(script, { + UDF1: [ + [1, 5], + [1, 5], + ], + }); + + expect(spans(result)).toEqual([ + [1, 5], + [6, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1"]); + }); + }); + + describe("malformed ranges", () => { + it("reads a reversed range as the span it describes", () => { + const result = segment(script, { UDF1: [[6, 2]] }); + + expect(spans(result)).toEqual([ + [1, 1], + [2, 6], + [7, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c2"]); + }); + + it("clamps a range that runs past the end of the file", () => { + const result = segment(script, { UDF1: [[8, 400]] }); + + expect(spans(result)).toEqual([ + [1, 7], + [8, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c2"]); + }); + + it("clamps a range that starts before the first line", () => { + const result = segment(script, { UDF1: [[-5, 3]] }); + + expect(spans(result)).toEqual([ + [1, 3], + [4, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1"]); + }); + + it("drops a range that lies entirely past the end and omits the UDF", () => { + const result = segment(script, { UDF1: [[40, 50]], UDF2: [[1, 10]] }); + + expect(spans(result)).toEqual([[1, 10]]); + expect(result.udfToCellUuids).toEqual({ UDF2: ["c1"] }); + }); + + it("ignores unusable entries but keeps the usable ones from the same UDF", () => { + const result = segment(script, { + UDF1: [[1, 4], "nonsense", null, [], [1, 2, 3], { start: 5 }, [5, 10]], + }); + + expect(spans(result)).toEqual([ + [1, 4], + [5, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1", "c2"]); + }); + + it("drops a range whose numeric bound is not finite", () => { + expect(segment(script, { UDF1: [[1, Infinity]] }).udfToCellUuids).toEqual({}); + expect(segment(script, { UDF1: [[NaN, 5]] }).udfToCellUuids).toEqual({}); + }); + + it("drops a range whose string bound is not a number", () => { + expect(segment(script, { UDF1: [["one", "four"]] }).udfToCellUuids).toEqual({}); + }); + + it("does not throw on entirely unusable input", () => { + expect(() => segment(script, { UDF1: "everything" })).not.toThrow(); + expect(() => segment(script, { UDF1: 42 })).not.toThrow(); + expect(segment(script, { UDF1: "everything" }).udfToCellUuids).toEqual({}); + }); + }); + + describe("accepted range shapes", () => { + it("accepts numeric strings", () => { + const result = segment(script, { UDF1: [["1", "4"]] }); + + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1"]); + expect(result.cells[0].endLine).toBe(4); + }); + + it("accepts the { start, end } object form", () => { + const result = segment(script, { UDF1: [{ start: 1, end: 4 }] }); + + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1"]); + expect(result.cells[0].endLine).toBe(4); + }); + + it("accepts a lone { start, end } object as a UDF's whole answer", () => { + const result = segment(script, { UDF1: { start: 1, end: 4 } }); + + expect(spans(result)).toEqual([ + [1, 4], + [5, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1"]); + }); + + it("accepts a bare [start, end] pair rather than a list of pairs", () => { + const result = segment(script, { UDF1: [1, 4] }); + + expect(spans(result)).toEqual([ + [1, 4], + [5, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1"]); + }); + + it("still reads a two-element list of pairs as two ranges, not one", () => { + const result = segment(script, { + UDF1: [ + [1, 2], + [9, 10], + ], + }); + + expect(spans(result)).toEqual([ + [1, 2], + [3, 8], + [9, 10], + ]); + expect(result.udfToCellUuids["UDF1"]).toEqual(["c1", "c3"]); + }); + + it("truncates fractional bounds", () => { + const result = segment(script, { UDF1: [[1.9, 4.2]] }); + + expect(spans(result)).toEqual([ + [1, 4], + [5, 10], + ]); + }); + }); + + describe("degenerate input", () => { + it("returns the whole file as one unmapped cell when no ranges are reported", () => { + const result = segment(script, {}); + + expect(spans(result)).toEqual([[1, 10]]); + expect(result.cells[0].source).toBe(script); + expect(result.udfToCellUuids).toEqual({}); + }); + + it("returns nothing for an empty script", () => { + expect(segment("", { UDF1: [[1, 3]] })).toEqual({ cells: [], udfToCellUuids: {} }); + }); + + it("returns nothing for a whitespace-only script", () => { + expect(segment("\n \n\t\n", { UDF1: [[1, 2]] }).cells).toEqual([]); + }); + + it("tolerates null and undefined ranges", () => { + expect(segment(script, null).udfToCellUuids).toEqual({}); + expect(spans(segment(script, undefined))).toEqual([[1, 10]]); + }); + + it("drops a blank-line-only gap rather than emitting an empty cell", () => { + const withBlankGap = ["import os", "", "", "print(os.getcwd())"].join("\n"); + const result = segment(withBlankGap, { UDF1: [[1, 1]], UDF2: [[4, 4]] }); + + expect(spans(result)).toEqual([ + [1, 1], + [4, 4], + ]); + expect(result.udfToCellUuids).toEqual({ UDF1: ["c1"], UDF2: ["c2"] }); + }); + }); + + describe("source fidelity", () => { + it("preserves blank lines and indentation inside a cell", () => { + const body = ["def f():", " x = 1", "", " return x"].join("\n"); + const result = segment(body, { UDF1: [[1, 4]] }); + + expect(result.cells[0].source).toBe(body); + }); + + it("normalizes CRLF and does not treat a trailing newline as a line", () => { + const result = segment("a\r\nb\r\n", { UDF1: [[1, 2]] }); + + expect(spans(result)).toEqual([[1, 2]]); + expect(result.cells[0].source).toBe("a\nb"); + }); + + it("keeps every line with content, each in exactly one cell", () => { + // A blank-only span between two reported ranges is dropped on purpose, so the cells do + // not reassemble to the input byte for byte. The guarantee that protects the user from + // losing code is narrower: no line with content is dropped, duplicated, or reordered. + const withBlankGap = ["import os", "", "", "x = compute()", "", "print(x)"].join("\n"); + const result = segment(withBlankGap, { UDF1: [[1, 1]], UDF2: [[4, 4]] }); + + const contentLines = (text: string) => text.split("\n").filter(line => line.trim() !== ""); + + expect(result.cells.flatMap(cell => contentLines(cell.source))).toEqual(contentLines(withBlankGap)); + }); + + it("gives every cell a distinct id from the supplied factory", () => { + const result = segment(script, { UDF1: [[1, 3]], UDF2: [[7, 9]] }); + const ids = result.cells.map(cell => cell.uuid); + + expect(ids).toEqual(["c1", "c2", "c3", "c4"]); + expect(new Set(ids).size).toBe(ids.length); + }); + + it("defaults to real uuids when no factory is supplied", () => { + const result = segmentScript(script, { UDF1: [[1, 10]] }); + + expect(result.cells[0].uuid).toMatch(/^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$/); + }); + }); +}); diff --git a/frontend/src/app/workspace/service/notebook-migration/script-segmentation.ts b/frontend/src/app/workspace/service/notebook-migration/script-segmentation.ts new file mode 100644 index 0000000000..b87c7040c9 --- /dev/null +++ b/frontend/src/app/workspace/service/notebook-migration/script-segmentation.ts @@ -0,0 +1,191 @@ +/** + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, + * software distributed under the License is distributed on an + * "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY + * KIND, either express or implied. See the License for the + * specific language governing permissions and limitations + * under the License. + */ + +import { v4 as uuidv4 } from "uuid"; + +/** + * Turns a Python script into notebook-style cells using the line ranges the LLM reported + * for each UDF. + * + * A notebook arrives already split into cells, and those cells are the join key for the + * cell<->operator mapping that drives highlighting. A script has no such boundaries, so the + * model is asked which lines it turned into which UDF, and the cells are derived from that + * answer here. + * + * The model's answer is untrusted: ranges can arrive reversed, overlapping, out of order, + * past the end of the file, or missing entirely. Every case is reconciled rather than + * rejected, because the ranges arrive alongside a workflow that already cost a full + * conversion, and a degraded mapping is worth more than a discarded result. + * + * Guarantees, given a non-empty script: + * - Every line with content lands in exactly one cell, whether or not a UDF claimed it. + * - Cells are disjoint and ordered by position in the file. + * - No cell straddles a reported boundary, so a cell claimed by a UDF is claimed whole. + * - Lines two UDFs both claim become one shared cell that maps to both, which is what + * `cell_to_operator` already expresses for notebooks. + */ + +// A 1-indexed, inclusive span of source lines. +interface LineRange { + start: number; + end: number; +} + +export interface DerivedCell { + uuid: string; + source: string; + // 1-indexed and inclusive, retained so callers can report or debug the split. + startLine: number; + endLine: number; +} + +export interface ScriptSegmentation { + cells: DerivedCell[]; + // UDF id -> the uuids of the cells it covers. A UDF whose ranges were all unusable is absent. + udfToCellUuids: Record<string, string[]>; +} + +/** + * Splits into lines, tolerating CRLF and a trailing newline (which is a terminator, not a line). + * + * Exported because the caller numbers these same lines when it builds the prompt. The numbers the + * model reports back are only meaningful if both sides agree on what counts as a line, so they must + * share one definition rather than each keep their own. + */ +export function splitScriptLines(source: string): string[] { + const lines = source.replace(/\r\n?/g, "\n").split("\n"); + if (lines.length > 0 && lines[lines.length - 1] === "") { + lines.pop(); + } + return lines; +} + +// Accepts a number or a numeric string, since models drift between the two. +function parseBound(value: unknown): number | null { + if (typeof value === "number") { + return Number.isFinite(value) ? Math.trunc(value) : null; + } + if (typeof value === "string" && value.trim() !== "") { + const parsed = Number(value); + return Number.isFinite(parsed) ? Math.trunc(parsed) : null; + } + return null; +} + +// Reads one range in either the [start, end] or { start, end } form. +function toLineRange(raw: unknown): LineRange | null { + let rawStart: unknown; + let rawEnd: unknown; + + if (Array.isArray(raw)) { + if (raw.length !== 2) return null; + [rawStart, rawEnd] = raw; + } else if (typeof raw === "object" && raw !== null) { + ({ start: rawStart, end: rawEnd } = raw as { start?: unknown; end?: unknown }); + } else { + return null; + } + + const start = parseBound(rawStart); + const end = parseBound(rawEnd); + if (start === null || end === null) return null; + + // A reversed range still names the span the model meant, so read it rather than drop it. + return start <= end ? { start, end } : { start: end, end: start }; +} + +// A UDF's ranges may arrive as a list of ranges, a single bare [start, end] pair, or one +// { start, end } object. Normalizes all three to a list before parsing. +function toRangeList(raw: unknown): unknown[] { + if (Array.isArray(raw)) { + const isBarePair = raw.length === 2 && raw.every(v => typeof v === "number" || typeof v === "string"); + return isBarePair ? [raw] : raw; + } + return typeof raw === "object" && raw !== null ? [raw] : []; +} + +function clampToSource(range: LineRange, lineCount: number): LineRange | null { + const start = Math.max(range.start, 1); + const end = Math.min(range.end, lineCount); + // Null when the range sat entirely past either end of the file. + return start <= end ? { start, end } : null; +} + +/** + * @param source the raw script contents + * @param rawRanges the model's reply, UDF id -> reported line ranges, unvalidated + * @param newUuid cell id factory, overridden by tests to keep output deterministic + */ +export function segmentScript( + source: string, + rawRanges: Record<string, unknown> | null | undefined, + newUuid: () => string = uuidv4 +): ScriptSegmentation { + const lines = splitScriptLines(source); + if (lines.length === 0) { + return { cells: [], udfToCellUuids: {} }; + } + + const udfRanges = new Map<string, LineRange[]>(); + for (const [udfId, raw] of Object.entries(rawRanges ?? {})) { + const ranges = toRangeList(raw) + .map(toLineRange) + .filter((range): range is LineRange => range !== null) + .map(range => clampToSource(range, lines.length)) + .filter((range): range is LineRange => range !== null); + if (ranges.length > 0) { + udfRanges.set(udfId, ranges); + } + } + + // Cut the file at every reported boundary. Slicing on the union of boundaries is what makes + // overlaps and gaps fall out on their own: the result is disjoint, covers every line, and no + // segment can span a boundary, so range membership below is an exact containment test. + const boundaries = new Set<number>([1, lines.length + 1]); + for (const ranges of udfRanges.values()) { + for (const { start, end } of ranges) { + boundaries.add(start); + boundaries.add(end + 1); + } + } + const cutPoints = [...boundaries].sort((a, b) => a - b); + + const cells: DerivedCell[] = []; + for (let i = 0; i < cutPoints.length - 1; i++) { + const startLine = cutPoints[i]; + const endLine = cutPoints[i + 1] - 1; + const text = lines.slice(startLine - 1, endLine).join("\n"); + // A span of only blank lines would render as an empty Jupyter cell, so it is dropped. + // Nothing is lost: every line carrying content still lands in a cell. + if (text.trim() === "") continue; + cells.push({ uuid: newUuid(), source: text, startLine, endLine }); + } + + const udfToCellUuids: Record<string, string[]> = {}; + for (const [udfId, ranges] of udfRanges) { + const uuids = cells + .filter(cell => ranges.some(range => cell.startLine >= range.start && cell.endLine <= range.end)) + .map(cell => cell.uuid); + if (uuids.length > 0) { + udfToCellUuids[udfId] = uuids; + } + } + + return { cells, udfToCellUuids }; +} diff --git a/frontend/src/assets/notebook_migration_tool/python_file_diagram.png b/frontend/src/assets/notebook_migration_tool/python_file_diagram.png new file mode 100644 index 0000000000..7b2081bc01 Binary files /dev/null and b/frontend/src/assets/notebook_migration_tool/python_file_diagram.png differ
