This is an automated email from the ASF dual-hosted git repository. github-merge-queue[bot] pushed a commit to branch gh-readonly-queue/main/pr-6016-964e915baae29601b8a81d0e90fe2ca7da35ceca in repository https://gitbox.apache.org/repos/asf/texera.git
commit 3da5f49c0a3727b1b57756265e84a5957551c036 Author: Mrudhulraj <[email protected]> AuthorDate: Tue Jul 21 20:21:52 2026 -0700 fix(query): Remove duplicated rows when datasets shared Publicly in the hub page (#6016) <!-- Thanks for sending a pull request (PR)! Here are some tips for you: 1. If this is your first time, please read our contributor guidelines: [Contributing to Texera](https://github.com/apache/texera/blob/main/CONTRIBUTING.md) 2. Ensure you have added or run the appropriate tests for your PR 3. If the PR is work in progress, mark it a draft on GitHub. 4. Please write your PR title to summarize what this PR proposes, we are following Conventional Commits style for PR titles as well. 5. Be sure to keep the PR description updated to reflect all changes. --> ### What changes were proposed in this PR? #### Issue - Duplicate datasets on hub landing page / hub search Symptom: A user creates a dataset, makes it public, and grants another user explicit access. When the grantee browses the hub, the dataset appears twice in the search results. Root cause: DatasetSearchQueryBuilder.constructFromClause produced this SQL: path: ```amber\src\main\scala\org\apache\texera\web\resource\dashboard\DatasetSearchQueryBuilder.scala:72``` ```sql SELECT DISTINCT ... FROM dataset LEFT JOIN dataset_user_access ON dua.did = dataset.did LEFT JOIN "user" ON ... WHERE (dua.uid = <ME>) OR (dataset.is_public = true) ``` For a dataset that is both public AND explicitly shared with the user, the LEFT JOIN produces one row per matching dataset_user_access row and the OR makes both branches true. **This applies similarly to worflows too.** #### Fix 1 — DatasetSearchQueryBuilder.constructFromClause Move the UID filter from the WHERE clause into the JOIN's ON clause so each dataset produces at most one joined row, and force the JOIN to FALSE when uid == null so the SELECT still references a valid table. ```scala val baseJoin = DATASET .leftJoin(DATASET_USER_ACCESS) .on(DATASET_USER_ACCESS.DID.eq(DATASET.DID)) .**and**(if (uid == null) DSL.**falseCondition**() else DATASET_USER_ACCESS.UID.eq(uid)) .leftJoin(USER) .on(USER.UID.eq(DATASET.OWNER_UID)) val condition: Condition = if (uid == null) { DATASET.IS_PUBLIC.eq(true) } else if (includePublic) { DATASET.IS_PUBLIC.eq(true).or(DATASET_USER_ACCESS.UID.isNotNull) } else { DATASET_USER_ACCESS.UID.isNotNull } baseJoin.where(condition) ``` Why `AND` `FALSE` for `uid == null`? The `SELECT` references `DATASET_USER_ACCESS.PRIVILEGE`. Without `dataset_user_access` in the `FROM`, DB throws missing FROM-clause entry for table "dataset_user_access". `AND` `FALSE` keeps the table in the FROM while making the JOIN yield NULL access columns — which is the correct semantic for "no explicit grant". Behavior matrix: | uid | includePublic | Matched datasets | |----------|---------------|---------------------------------------------| | null | (n/a) | Public only | | not null | false | Datasets with explicit access of logged-in user only | | not null | true | Public + logged-in explicit access **(no duplicates)** | <!-- Please clarify what changes you are proposing. The purpose of this section is to outline the changes. Here are some tips for you: 1. If you propose a new API, clarify the use case for a new API. 2. If you fix a bug, you can clarify why it is a bug. 3. If it is a refactoring, clarify what has been changed. 3. It would be helpful to include a before-and-after comparison using screenshots or GIFs. 4. Please consider writing useful notes for better and faster reviews. --> ### Any related issues, documentation, discussions? Fixes #5957 <!-- Please use this section to link other resources if not mentioned already. 1. If this PR fixes an issue, please include `Fixes #1234`, `Resolves #1234` or `Closes #1234`. If it is only related, simply mention the issue number. 2. If there is design documentation, please add the link. 3. If there is a discussion in the mailing list, please add the link. --> ### How was this PR tested? Testing: 1. Created a new DatasetResourcespec for basic unit-testing. 2. Manual: Using database checks and UI workflow testing. Share dataset to user with READ/WRITE permissions + Publicly: <img width="682" height="795" alt="Share Permissions" src="https://github.com/user-attachments/assets/ad42ae50-0153-4aec-bb72-4b7b084f4c91" /> Observation (before fix): Two datasets are listed for the shared user (User access permission + Public access (NONE)) <img width="1900" height="895" alt="testUser perspective" src="https://github.com/user-attachments/assets/69ce507b-fdbc-4c87-99d0-e7fa907fbf39" /> Result (after fix): 1 dataset of the user access permission <img width="1915" height="717" alt="testUser result 1" src="https://github.com/user-attachments/assets/2fd0a06e-3e5e-4a9e-b550-0245619b42f8" /> <!-- If tests were added, say they were added here. Or simply mention that if the PR is tested with existing test cases. Make sure to include/update test cases that check the changes thoroughly including negative and positive cases if possible. If it was tested in a way different from regular unit tests, please clarify how you tested step by step, ideally copy and paste-able, so that other reviewers can test and check, and descendants can verify in the future. If tests were not added, please describe why they were not added and/or why it was difficult to add. --> ### Was this PR authored or co-authored using generative AI tooling? AI tools used for generating the spec files. <!-- If generative AI tooling has been used in the process of authoring this PR, please include the phrase: 'Generated-by: ' followed by the name of the tool and its version. If no, write 'No'. Please refer to the [ASF Generative Tooling Guidance](https://www.apache.org/legal/generative-tooling.html) for details. --> --- .../dashboard/DatasetSearchQueryBuilder.scala | 47 +++-- .../dashboard/file/DatasetResourceSpec.scala | 228 +++++++++++++++++++++ 2 files changed, 258 insertions(+), 17 deletions(-) diff --git a/amber/src/main/scala/org/apache/texera/web/resource/dashboard/DatasetSearchQueryBuilder.scala b/amber/src/main/scala/org/apache/texera/web/resource/dashboard/DatasetSearchQueryBuilder.scala index 64c8c31106..0cda3eecdc 100644 --- a/amber/src/main/scala/org/apache/texera/web/resource/dashboard/DatasetSearchQueryBuilder.scala +++ b/amber/src/main/scala/org/apache/texera/web/resource/dashboard/DatasetSearchQueryBuilder.scala @@ -66,30 +66,38 @@ object DatasetSearchQueryBuilder extends SearchQueryBuilder with LazyLogging { params: DashboardResource.SearchQueryParams, includePublic: Boolean = false ): TableLike[_] = { + // Case 1: if `uid` is (set) and `includePublic` is false + // -> return ONLY datasets that given `uid` has explicit access to. + // Case 2: if `uid` is (null) and `includePublic` is true + // -> return ONLY datasets that are public + // Case 3: if `uid` is (set) and `includePublic` is true + // -> Union of datasets that are public and explicitly shared with user is returned + // Case 4: if `uid` is (null) and `includePublic` is false + // -> return public datasets by default as user might not be logged in val baseJoin = DATASET .leftJoin(DATASET_USER_ACCESS) .on(DATASET_USER_ACCESS.DID.eq(DATASET.DID)) + .and(if (uid == null) DSL.falseCondition() else DATASET_USER_ACCESS.UID.eq(uid)) .leftJoin(USER) .on(USER.UID.eq(DATASET.OWNER_UID)) - // Default condition starts as true, ensuring all datasets are selected initially. - var condition: Condition = DSL.trueCondition() - - if (uid == null) { - // If `uid` is null, the user is not logged in or performing a public search - // We only select datasets marked as public - condition = DATASET.IS_PUBLIC.eq(true) - } else { - // When `uid` is present, we add a condition to only include datasets with direct user access. - val userAccessCondition = DATASET_USER_ACCESS.UID.eq(uid) - - if (includePublic) { - // If `includePublic` is true, we extend visibility to public datasets as well. - condition = userAccessCondition.or(DATASET.IS_PUBLIC.eq(true)) + // Set the `condition` where clause here + val condition: Condition = + if (uid == null) { + // Case 2 and 4 + // Get all the public datasets by default + DATASET.IS_PUBLIC.eq(true) } else { - condition = userAccessCondition + if (includePublic) { + // Case 3 + // Get all the datasets that `uid` has access to and the public datasets + DATASET.IS_PUBLIC.eq(true).or(DATASET_USER_ACCESS.UID.isNotNull) + } else { + // Case 1 + // If `includePublic` is false get only user accessible datasets + DATASET_USER_ACCESS.UID.isNotNull + } } - } baseJoin.where(condition) } @@ -140,7 +148,12 @@ object DatasetSearchQueryBuilder extends SearchQueryBuilder with LazyLogging { val dd = DashboardDataset( dataset, owner.getEmail, - record.get(DATASET_USER_ACCESS.PRIVILEGE, classOf[PrivilegeEnum]), + Option( + record.get( + DATASET_USER_ACCESS.PRIVILEGE, + classOf[PrivilegeEnum] + ) + ).getOrElse(PrivilegeEnum.NONE), dataset.getOwnerUid == uid, size ) diff --git a/amber/src/test/scala/org/apache/texera/web/resource/dashboard/file/DatasetResourceSpec.scala b/amber/src/test/scala/org/apache/texera/web/resource/dashboard/file/DatasetResourceSpec.scala new file mode 100644 index 0000000000..1d4b5635e0 --- /dev/null +++ b/amber/src/test/scala/org/apache/texera/web/resource/dashboard/file/DatasetResourceSpec.scala @@ -0,0 +1,228 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, + * software distributed under the License is distributed on an + * "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY + * KIND, either express or implied. See the License for the + * specific language governing permissions and limitations + * under the License. + */ + +package org.apache.texera.web.resource.dashboard.file + +import org.apache.texera.auth.SessionUser +import org.apache.texera.dao.MockTexeraDB +import org.apache.texera.dao.jooq.generated.enums.UserRoleEnum +import org.apache.texera.dao.jooq.generated.tables.pojos.User +import org.apache.texera.web.resource.dashboard.DashboardResource.SearchQueryParams +import org.apache.texera.web.resource.dashboard.user.dataset.DatasetResource.DashboardDataset +import org.apache.texera.web.resource.dashboard.{FulltextSearchQueryUtils} +import org.apache.texera.web.resource.dashboard.DatasetSearchQueryBuilder +import org.scalatest.flatspec.AnyFlatSpec +import org.apache.texera.dao.jooq.generated.tables.daos.{UserDao, DatasetDao, DatasetUserAccessDao} +import org.apache.texera.dao.jooq.generated.enums.PrivilegeEnum +import org.apache.texera.dao.jooq.generated.tables.pojos.{Dataset, DatasetUserAccess} +import org.scalatest.{BeforeAndAfterAll, BeforeAndAfterEach} +import java.time.OffsetDateTime +import java.util +import org.apache.texera.web.resource.dashboard.SearchQueryBuilder.DATASET_RESOURCE_TYPE + +class DatasetResourceSpec + extends AnyFlatSpec + with BeforeAndAfterAll + with BeforeAndAfterEach + with MockTexeraDB { + + // An example creation time to test Account Creation Time attribute + private val exampleCreationTime: OffsetDateTime = + OffsetDateTime.parse("2025-01-01T00:00:00Z") + + private val ownerUser: User = { + val user = new User + user.setUid(Integer.valueOf(1)) + user.setName("owner_user") + user.setRole(UserRoleEnum.ADMIN) + user.setEmail("[email protected]") + user.setPassword("123") + user.setComment("test_comment") + user.setAccountCreationTime(exampleCreationTime) + user + } + + private val testUser: User = { + val user = new User + user.setUid(Integer.valueOf(2)) + user.setName("test_user") + user.setEmail("[email protected]") + user.setRole(UserRoleEnum.REGULAR) + user.setPassword("123") + user.setComment("test_comment2") + user.setAccountCreationTime(exampleCreationTime) + user + } + + private val testDatasetRecord: Dataset = { + val dataset = new Dataset() + dataset.setName("test_dataset1") + dataset.setDescription("keyword_in_dataset_description") + dataset.setIsPublic(true) + dataset.setDid(Integer.valueOf(1)) + dataset + } + + private val sessionUser1: SessionUser = { + new SessionUser(ownerUser) + } + + private val sessionUser2: SessionUser = { + new SessionUser(testUser) + } + + // get context lazily + private lazy val datasetDao: DatasetDao = { + new DatasetDao(getDSLContext.configuration()) + } + + private lazy val datasetUserAccessDao: DatasetUserAccessDao = { + new DatasetUserAccessDao(getDSLContext.configuration()) + } + + override protected def beforeAll(): Unit = { + initializeDBAndReplaceDSLContext() + FulltextSearchQueryUtils.usePgroonga = false // disable pgroonga + // add test user directly + val userDao = new UserDao(getDSLContext.configuration()) + userDao.insert(ownerUser) + userDao.insert(testUser) + } + + override protected def beforeEach(): Unit = { + // Clean up environment before each test case + } + + override protected def afterEach(): Unit = { + // 1. Delete access rows before the dataset + val datasetUserAccessDao = new DatasetUserAccessDao(getDSLContext.configuration()) + getDSLContext + .deleteFrom(org.apache.texera.dao.jooq.generated.tables.DatasetUserAccess.DATASET_USER_ACCESS) + .execute() + // 2. Fetch all datasets owned by the owner + val datasets = datasetDao.fetchByOwnerUid(ownerUser.getUid()) + if (!datasets.isEmpty) { + datasetDao.delete(datasets) + } + } + + override protected def afterAll(): Unit = { + shutdownDB() + } + + private def getKeywordsArray(keywords: String*): util.ArrayList[String] = { + val keywordsList = new util.ArrayList[String]() + for (keyword <- keywords) { + keywordsList.add(keyword) + } + keywordsList + } + + private def assertSameDataset(a: Dataset, b: DashboardDataset): Unit = { + assert(a.getName == b.dataset.getName) + } + + "User.accountCreationTime" should "be persisted and retrievable via UserDao" in { + val userDao = new UserDao(getDSLContext.configuration()) + val u1 = userDao.fetchOneByUid(Integer.valueOf(1)) + val u2 = userDao.fetchOneByUid(Integer.valueOf(2)) + + assert(u1.getAccountCreationTime != null) + assert(u2.getAccountCreationTime != null) + + assert(u1.getAccountCreationTime.isEqual(exampleCreationTime)) + assert(u2.getAccountCreationTime.isEqual(exampleCreationTime)) + } + + it should "remain unchanged when updating unrelated fields" in { + val userDao = new UserDao(getDSLContext.configuration()) + val u1 = userDao.fetchOneByUid(Integer.valueOf(1)) + val originalTime = u1.getAccountCreationTime + + u1.setComment("updated_comment") + userDao.update(u1) + + val test_u1 = userDao.fetchOneByUid(Integer.valueOf(1)) + assert(test_u1.getAccountCreationTime.isEqual(originalTime)) + } + + "DatasetResource /owner view" should "get deduplicated datasets created by owner and shared publicly" in { + // Only metadatas of dataset and the user is maintained - no dataset is actually created in LakeFS + // Create dataset + val datasetDao = new DatasetDao(getDSLContext.configuration()) + testDatasetRecord.setOwnerUid(ownerUser.getUid) + datasetDao.insert(testDatasetRecord) + + // Give write access to the owner user + datasetUserAccessDao.insert( + new DatasetUserAccess( + testDatasetRecord.getDid, + ownerUser.getUid, + PrivilegeEnum.WRITE + ) + ) + + // Build the query - bypasses DashboarResource + val query = + DatasetSearchQueryBuilder.constructQuery( + ownerUser.getUid, + SearchQueryParams(resourceType = DATASET_RESOURCE_TYPE), + includePublic = true + ) + // Assert the length of returned dataset + val datasetEntryList = getDSLContext.fetch(query) + assert(datasetEntryList.size() == 1) + } + + "/search API" should "deduplicate datasets shared both publicly and explicitly with NON-WRITE permissions" in { + // Only metadatas of dataset and the user is maintained - no dataset is actually created in LakeFS + // Create dataset + val datasetDao = new DatasetDao(getDSLContext.configuration()) + testDatasetRecord.setOwnerUid(ownerUser.getUid) + datasetDao.insert(testDatasetRecord) + + // Give write access to the owner user + datasetUserAccessDao.insert( + new DatasetUserAccess( + ownerUser.getUid, + testDatasetRecord.getDid, + PrivilegeEnum.WRITE + ) + ) + + datasetUserAccessDao.insert( + new DatasetUserAccess( + testDatasetRecord.getDid, + testUser.getUid, + PrivilegeEnum.READ + ) + ) + + // Build the query + val query = + DatasetSearchQueryBuilder.constructQuery( + testUser.getUid, + SearchQueryParams(resourceType = DATASET_RESOURCE_TYPE), + includePublic = true + ) + // Assert the length of returned dataset + val datasetEntryList = getDSLContext.fetch(query) + assert(datasetEntryList.size() == 1) + } +}
