[ 
https://issues.apache.org/jira/browse/PHOENIX-7032?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17782397#comment-17782397
 ] 

ASF GitHub Bot commented on PHOENIX-7032:
-----------------------------------------

kadirozde commented on PR #1701:
URL: https://github.com/apache/phoenix/pull/1701#issuecomment-1791860357

   > The following index does not generate any index rows -
   > 
   > CREATE TABLE IF NOT EXISTS S.T1(OID CHAR(15) NOT NULL,KP CHAR(3) NOT NULL, 
ID1 INTEGER, ID2 VARCHAR, ID3 INTEGER, ROW_ID CHAR(15), COL1 VARCHAR, COL2 
VARCHAR, CREATED_DATE DATE,CREATED_BY CHAR(15),LAST_UPDATE DATE,LAST_UPDATE_BY 
CHAR(15),SYSTEM_MODSTAMP DATE CONSTRAINT pk PRIMARY KEY (OID,KP)) 
MULTI_TENANT=true,COLUMN_ENCODED_BYTES=0;
   > 
   > CREATE INDEX IF NOT EXISTS P_T1_COL1_2_NOT_NULL_INDEX ON S.T1 (KP, COL1, 
COL2) INCLUDE (ID1, ID2, ID3) WHERE COL1 IS NOT NULL AND COL2 IS NOT NULL;
   > 
   > UPSERT INTO S.T1(OID, KP, CREATED_DATE, ID1, ID2, ID3, ROW_ID, COL1, COL2, 
SYSTEM_MODSTAMP) VALUES('00D0x0000000001', '001', now(), 12200, 'id1-0001', 
14566, '00R0r0000000001', 'col1-0001', 'col2-0001', now());
   > 
   > Similar behavior for the following index CREATE INDEX IF NOT EXISTS 
P_T1_ID2_COL1_NOT_NULL_INDEX ON S.T1 (KP, COL1, COL2) INCLUDE (ID1, ID2, ID3) 
WHERE ID2 = 'id1-0001' AND COL1 IS NOT NULL;
   
   Great catch @jpisaac ! The list of cells are searched using a binary search 
during expression evaluation and thus they need to be sorted before evaluated. 
I have just fixed it.




> Partial Global Secondary Indexes
> --------------------------------
>
>                 Key: PHOENIX-7032
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-7032
>             Project: Phoenix
>          Issue Type: New Feature
>            Reporter: Kadir Ozdemir
>            Assignee: Kadir Ozdemir
>            Priority: Major
>
> The secondary indexes supported in Phoenix have been full indexes such that 
> for every data table row there is an index row. Generating an index row for 
> every data table row is not always required. For example, some use cases do 
> not require index rows for the data table rows in which indexed column values 
> are null. Such indexes are called sparse indexes. Partial indexes generalize 
> the concept of sparse indexing and allow users to specify the subset of the 
> data table rows for which index rows will be maintained. This subset is 
> specified using a WHERE clause added to the CREATE INDEX DDL statement.
> Partial secondary indexes were first proposed by Michael Stonebraker 
> [here|https://dsf.berkeley.edu/papers/ERL-M89-17.pdf]. Since then several SQL 
> databases (e.g., 
> [Postgres|https://www.postgresql.org/docs/current/indexes-partial.html] and 
> [SQLite|https://www.sqlite.org/partialindex.html])  and NoSQL databases 
> (e.g., [MongoDB|https://www.mongodb.com/docs/manual/core/index-partial/]) 
> have supported some form of partial indexes. It is challenging to allow 
> arbitrary WHERE clauses in DDL statements. For example, Postgres does not 
> allow subqueries in these where clauses and SQLite supports much more 
> restrictive where clauses. 
> Supporting arbitrary where clauses creates challenges for query optimizers in 
> deciding the usability of a partial index for a given query. If the set of 
> data table rows that satisfy the query is a subset of the data table rows 
> that the partial index points back, then the query can use the index. Thus, 
> the query optimizer has to decide if the WHERE clause of the query implies 
> the WHERE clause of the index. 
> Michael Stonebraker [here|https://dsf.berkeley.edu/papers/ERL-M89-17.pdf] 
> suggests that an index WHERE clause is a conjunct of simple terms, i.e: 
> i-clause-1 and i-clause-2 and ... and i-clause-m where each clause is of the 
> form <column> <operator> <constant>. Hence, the qualification can be 
> evaluated for each tuple in the indicated relation without consulting 
> additional tuples. 
> Phoenix partial indexes will initially support a more general set of index 
> WHERE clauses that can be evaluated on a single row with the following 
> exceptions
>  * Subqueries are not allowed.
>  * Like expressions are allowed with very limited support such that an index 
> WHERE clause with like expressions can imply/contain a query if the query has 
> the same like expressions that the index WHERE clause has.
>  * Comparison between columns are allowed without supporting transitivity, 
> for example, a > b and b > c does not imply a > c.
> Partial indexes will be supported initially for global secondary indexes, 
> i.e.,  covered global indexes and uncovered global indexes. The local 
> secondary indexes will be supported in future.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to