influxdb

Commit Graph

Author	SHA1	Message	Date
wayne	0fffcc8c37	refactor: introduce influxdb3_types crate (#25946 ) Partially fixes https://github.com/influxdata/influxdb/issues/24672 * move most HTTP req/resp types into `influxdb3_types` crate * removes the use of locally-scoped request type structs from the `influxdb3_client` crate * fix plugin dependency/package install bug * it looks like the `DELETE` http method was being used where `POST` was expected for `/api/v3/configure/plugin_environment/install_packages` and `/api/v3/configure/plugin_environment/install_requirements`	2025-02-03 11:28:47 -07:00
Michael Gattozzi	20fdc7b51b	feat: stream data back for CSV and JSON queries (#25927 ) This commit allows us to stream data back for CSV and JSON formatted queries. Prior to this we would buffer up all of the data in memory before sending it back. Now we can make it so that we only buffer in one RecordBatch at a time to reduce memory overhead. Note that due to the way the APIs for writers work and for how Body in hyper 0.14 works we can't use a streaming body that we can write too. This in turn means we have to use a manually written Future state machine, which works but is far from ideal. Note this does not include the pretty and parquet files as streamable. I'm attempting to get the pretty one to be streamable, but I don't think that this one and parquet are as likely to be streamable back to the user. In general we might want to discourage these formats from being used.	2025-01-31 13:25:30 -05:00
wayne	05da40fa9b	fix: clarify table creation conflict error message (#25936 ) Also include a basic CLI integration test to exemplify the new error message.	2025-01-30 13:08:35 -07:00
Jackson Newhouse	8840d99e9d	feat(processing_engine): integration with virtual environments. (#25895 ) * feat(processing_engine): integration with virtual environments. * feat: Initial scaffolding for environment managers (pip, pipx, uv). * feat(processing_engine): CLI for package management, remove pipx support. * feat(processing_engine): test installations in virtualenvs. * feat(processing_engine): Automatically setup virtual environment on startup.	2025-01-28 15:30:17 -08:00
Paul Dix	9a5424693c	feat: Update trigger creation to validate plugin file present (#25918 ) This updates trigger creation to load the plugin file before creating the trigger. Another small change is to make Github references use filenames and paths identical to what they would be in the plugin-dir. This makes it a little easier to have the plugins repo local and develop against it and then be able to reference the same file later with gh: once it's up on the repo.	2025-01-27 12:31:49 -05:00
Paul Dix	d49276a7fb	feat: Refactor plugins to only require creating trigger (#25914 ) This refactors plugins and triggers so that plugins no longer need to be "created". Since plugins exist in either the configured local directory or on the Github repo, a user now only needs to create a trigger and reference the plugin filename. Closes #25876	2025-01-27 11:26:46 -05:00
Michael Gattozzi	43e186d761	feat: add no_sync write_lp param for fast writes (#25902 )	2025-01-24 13:34:38 -05:00
praveen-influx	4ef972eab4	feat: first stab at locally updating parquet cache (#25904 ) * feat: first stab at locally updating parquet cache closes: https://github.com/influxdata/influxdb/issues/25887 * refactor: use enums to separate out the modes This commit introduced the `Immediate` and `Eventual` modes for fulfilling the cache request. In immediate mode since the data is readily available to be cached, we can avoid extra requests to object store. part of: https://github.com/influxdata/influxdb/issues/25887	2025-01-24 14:36:06 +00:00
Trevor Hilton	d451ef0de6	refactor: writer-id to node-id (#25905 )	2025-01-23 18:09:24 -05:00
Michael Gattozzi	63bd5096f5	feat: loosen 72 hour query/write restriction (#25890 ) This commit does a few key things: - Removes the 72 hour query and write restrictions in Core - Limits the queries to a default number of parquet files. We chose 432 as this is about 72 hours using default settings for the gen1 timeblock - The file limit can be increased, but the help text and error message when exceeded note that query performance will likely be degraded as a result. - We warn users to use smaller time ranges if possible if they hit this query error With this we eliminate the hard restriction we have in place, but instead create a soft one that users can choose to take the performance hit with. If they can't take that hit then it's recomended that they upgrade to Enterprise which has the compactor built in to make performant historical queries.	2025-01-23 10:02:26 -05:00
Trevor Hilton	44ca7a4d36	refactor: reduce catalog locks when getting chunks (#25896 ) * refactor: reduce catalog locks when getting chunks The main refactor was to change the ChunkContainer trait to use the DatabaseSchema and TableDefinition types directly in the signature, vs. the names, which then required an additional catalog lock and lookups for both entities. This was already handled upstream in the QueryTable, so there was no need to do the lookups again. This required the addition of a test helper in influxdb3_write::test_helpers that provides convenience methods for getting record batches from the WriteBuffer. We have been implementing such a method manually in several places, so this is nice to have it unified. This provides a blanket impl so that anything implementing WriteBuffer gets the method. Some other house cleaning was included. * refactor: clean up test helpers in influxdb3_write * refactor: pass original df filters forward with ChunkFilter * chore: clippy	2025-01-22 14:38:46 -05:00
Paul Dix	e87f8862b3	feat: Add request plugin capability (#25864 ) * feat: Add request plugin capability Adds the request plugin type. Triggers can be bound to an API endpoint at /api/v3/engine/<path>. Requests will get yielded to the plugin with the query parameters, request parameters, and request body. I didn't implement the test endpoint for this plugin type as it seems much more natural for users to save the file and make a new request. Once #25863 is done it'll make it very easy. Closes #25862 * chore: fix spelling in error message	2025-01-21 20:22:27 -08:00
praveen-influx	447f66d9a7	fix: allow default browser header (#25885 ) Although the `format` in the request is used, the value coming through the header is parsed earlier. So, when that lookup in the header fails an error is returned (`InvalidMimeType`). In this commit, there are extra checks to allow the default `Accept` header values that come from the browser by defaulting it to `json` closes: https://github.com/influxdata/influxdb/issues/25874	2025-01-21 13:43:07 +00:00
Jackson Newhouse	1d8d3d66fc	feat(processing_engine): Add cron plugins and triggers to the processing engine. (#25852 ) * feat(processing_engine): Add cron plugins and triggers to the processing engine. * feat(processing_engine): switch from 'cron plugin' to 'schedule plugin', use TimeProvider. * feat(processing_engine): add test for test scheduled plugin.	2025-01-18 07:18:18 -05:00
Trevor Hilton	8ea6d65f8e	chore: bring back enterprise changes (#25858 )	2025-01-17 13:24:32 -05:00
Jackson Newhouse	244ac567a8	fix(processing_engine): start triggers when server is started. (#25848 )	2025-01-16 16:39:42 -08:00
Trevor Hilton	b8a94488b5	feat: support v1 query API GROUP BY semantics (#25845 ) This updates the v1 /query API hanlder to handle InfluxDB v1's unique query response structure when GROUP BY clauses are provided. The distinction is in the addition of a "tags" field to the emitted series data that contains a map of the GROUP BY tags along with their distinct values associated with the data in the "values" field. This required splitting the QueryExecutor into two query paths for InfluxQL and SQL, as this allowed for handling InfluxQL query parsing in advance of query planning. A set of snapshot tests were added to check that it all works.	2025-01-16 11:59:01 -05:00
Trevor Hilton	db24a62658	refactor: change host-id to writer-id (#25804 ) This changes the CLI arg `host-id` to `writer-id` to more accurately indicate meaning. This changes also goes through the codebase and changes struct fields, methods, and variables to use the term `writer_id` or `writer_identifier_prefix` instead of `host_id` etc., to make the meaning clear in the code. This also changes the catalog serialization to use the field `writer_id` instead of `host_id`, which is breaking change.	2025-01-12 11:40:47 -05:00
Paul Dix	491a37b0d4	feat: Update create plugin to use server file (#25803 ) This updates the create plugin API and CLI so that it doesn't take the pugin code, but instead takes a file name of a file that must be in the plugin-dir of the server. It returns an error if the plugin-dir is not configured or if the file isn't there. Also updates the WAL and catalog so that it doesn't store the plugin code directly. The code is read from disk one time when the plugin runs. Closes #25797	2025-01-11 21:02:51 -05:00
praveen-influx	50963443a4	feat: introduce num wal files to keep (#25801 ) * feat: introduce num wal files to keep This commit allows a configurable number of wal files to be left behind in OS. This is necessary as enterprise replicas rely on these files. closes: https://github.com/influxdata/influxdb/issues/25788 * refactor: address PR feedback * refactor: address PR comment	2025-01-12 00:33:13 +00:00
Trevor Hilton	0bdc2fa953	chore: patch enterprise back to core (#25798 )	2025-01-11 17:26:41 -05:00
Paul Dix	e421baf0bc	feat: include specific plugin error to http response (#25794 )	2025-01-11 14:32:36 -05:00
Paul Dix	bdcdb4f296	refactor: rename plugin activate/deactivate to enable/disable (#25793 ) Closes #25789	2025-01-11 13:40:24 -05:00
Paul Dix	e8422a240a	feat: Wire up arguments to wal plugin trigger (#25783 ) This allows the user to specify arguments that will be passed to each execution of a wal plugin trigger. The CLI test was updated to check this end to end. Closes #25655	2025-01-10 16:58:18 -05:00
Jackson Newhouse	954416d4d9	fix(processing_engine): enable downstream system-py feature (#25784 )	2025-01-10 13:44:24 -08:00
Paul Dix	0da0785960	feat: Finish wiring up WAL plugin with trigger (#25781 ) This updates the WAL so that it can have new file notifiers added that will get updated when the wal flushes. The processing engine now implements the WALNotifier trait. I've updated the CLI test for creating a trigger to run and end-to-end test that defines a plugin, creates a trigger, writes data into the database, triggering the plugin, which writes summary statistics back into the database in a different table. The test queries the destination table to confirm that the plugin worked.	2025-01-10 12:56:49 -05:00
Trevor Hilton	c71dafc313	refactor: rename metadata cache to distinct value cache (#25775 )	2025-01-10 08:48:51 -05:00
Paul Dix	7230148b58	feat: Update WAL plugin for new structure (#25777 ) * feat: Update WAL plugin for new structure This ended up being a very large change set. In order to get around circular dependencies, the processing engine had to be moved into its own crate, which I think is ultimately much cleaner. Unfortunately, this required changing a ton of things. There's more testing and things to add on to this, but I think it's important to get this through and build on it. Importantly, the processing engine no longer resides inside the write buffer. Instead, it is attached to the HTTP server. It is now able to take a query executor, write buffer, and WAL so that the full range of functionality of the server can be exposed to the plugin API. There are a bunch of system-py feature flags littered everywhere, which I'm hoping we can remove soon. * refactor: PR feedback	2025-01-10 05:52:33 -05:00
Paul Dix	2d18a61949	feat: Add query API to Python plugins (#25766 ) This ended up being a couple things rolled into one. In order to add a query API to the Python plugin, I had to pull the QueryExecutor trait out of server into a place so that the python crate could use it. This implements the query API, but also fixes up the WAL plugin test CLI a bit. I've added a test in the CLI section so that it shows end-to-end operation of the WAL plugin test API and exercise of the entire Plugin API. Closes #25757	2025-01-09 20:13:20 -05:00
Trevor Hilton	23866ef3d1	fix: /query API emits empty result as JSON (#25769 )	2025-01-09 09:26:30 -05:00
praveen-influx	6e2e39cd4c	feat: snapshot when wal buffer is empty (#25765 ) * feat: snapshot when wal buffer is empty - This commit changes the functionality to allow snapshots to happen even when wal buffer is empty. For snapshots wal periods are still required but not the wal buffer. To allow this, we write a no-op into wal file with snapshot details. This enables force snapshotting functionality closes: https://github.com/influxdata/influxdb/issues/25685 * refactor: address PR feedback	2025-01-09 12:12:37 +00:00
Trevor Hilton	dfc853d903	feat: handle params in request body for /query API (#25762 ) Closes #25749 This changes the `/query` API handler so that the parameters can be passed in either the request URI or in the request body for either a `GET` or `POST` request. Parameters can be specified in the URI, the body, or both; if they are specified in both places, those in the body will take precedent. Error variants in the HTTP server code related to missing request parameters were updated to return `400` status.	2025-01-07 19:53:29 -05:00
Trevor Hilton	3efd49066e	fix: bind to correct port for e2e tests (#25758 ) * fix: bind to correct port for e2e tests This also fixes up some log messages on server start for naming * chore: do notpass value in TEST_LOG env var to CI tests	2025-01-07 12:23:59 -05:00
Trevor Hilton	6524f383ba	feat: show databases CLI/API (#25748 ) _Follows #25737 (keeping in draft until that merges)_ Closes #25745 This PR provides both a CLI and underlying API for listing databases in the InfluxDB 3 Core server. Details are below. There was already a method to list databases for the query executor for InfluxQL; this works by exposing that via the `HttpApi` in `influxdb3_server`. However, one thing that we may address is that the query result for that uses `iox::database` as the column name. If we are removing references to `iox`, then we may want to just have it as `database`. I left it as is, for now, because I wanted to keep code churn down and wasn't sure why we use that prefix in the first place for the `SHOW DATABASES` and `SHOW RETENTION POLICIES` InfluxQL queries. ## Details ### CLI This PR provides the `influxdb3 show` CLI: ``` influxdb3 show -h List resources on the InfluxDB 3 Core server Usage: influxdb3 show <COMMAND> Commands: databases List databases help Print this message or the help of the given subcommand(s) Options: -h, --help Print help information ``` with the ability to list databases: ``` influxdb3 show databases -h List databases Usage: influxdb3 show databases [OPTIONS] Options: -H, --host <HOST_URL> The host URL of the running InfluxDB 3 Core server [env: INFLUXDB3_HOST_URL=] [default: http://127.0.0.1:8181] --token <AUTH_TOKEN> The token for authentication with the InfluxDB 3 Core server [env: INFLUXDB3_AUTH_TOKEN=] --show-deleted Include databases that were marked as deleted in the output --format <OUTPUT_FORMAT> The format in which to output the list of databases [default: pretty] [possible values: pretty, json, json_lines, csv] -h, --help Print help information ``` Since this uses the query executor, we can pass a `--format` argument to get the output as JSON, CSV, or JSONL, but by default, it uses the `pretty` format: ``` influxdb3 show databases +---------------+ \| iox::database \| +---------------+ \| bar \| +---------------+ ``` The `--show-deleted` flag will have the `deleted` column displayed as well as any databases that have been marked as deleted: ``` influxdb3 show databases --show-deleted +---------------------+---------+ \| iox::database \| deleted \| +---------------------+---------+ \| bar \| false \| \| foo-20250105T202949 \| true \| +---------------------+---------+ ``` ### API The API to list databases can be invoked via: ``` GET /api/v3/configure/database ``` with optional parameters: * `format`: `pretty`, `json`, `csv`, `parquet`, or `jsonl` * `show_deleted`: `bool`, defaults to `false` Note that `database` is singular in the API endpoint, to be consistent with the other database related create/delete API endpoints. We could change it to be plural `databases` if that is the convention we want to go with.	2025-01-06 21:08:12 -05:00
Michael Gattozzi	f793d31f63	feat: Cleanup CLI flags for InfluxDB 3 Core (#25737 ) This makes quite a few major changes to our CLI and how users interact with it: 1. All commands are now in the form <verb> <noun> this was to make the commands consistent. We had last-cache as a noun, but serve as a verb in the top level. Given that we could only create or delete All noun based commands have been move under a create and delete command 2. --host short form is now -H not -h which is reassigned to -h/--help for shorter help text and is in line with what users would expect for a CLI 3. Only the needed items from clap_blocks have been moved into `influxdb3_clap_blocks` and any IOx specific references were changed to InfluxDB 3 specific ones 4. References to InfluxDB 3.0 OSS have been changed to InfluxDB 3 Core in our CLI tools 5. --dbname has been changed to --database to be consistent with --table in many commands. The short -d flag still remains. In the create/ delete command for the database however the name of the database is a positional arg e.g. `influxbd3 create database foo` rather than `influxdb3 database create --dbname foo` 6. --table has been removed from the delete/create command for tables and is now a positional arg much like database 7. clap_blocks was removed as dependency to avoid having IOx specific env vars 8. --cache-name is now an optional positional arg for last_cache and meta_cache 9. last-cache/meta-cache commands are now last_cache and meta_cache respectively Unfortunately we have quite a few options to run the software and I couldn't cut down on them, but at least with this commands and options will be more discoverable and we have full control over our CLI options now. Closes #25646	2025-01-06 18:51:55 -05:00
Paul Dix	1ce6a24c3f	feat: Implement WAL plugin test API (#25704 ) * feat: Implement WAL plugin test API This implements the WAL plugin test API. It also introduces a new API for the Python plugins to be called, get their data, and call back into the database server. There are some things that I'll want to address in follow on work: * CLI tests, but will wait on #25737 to land for a refactor of the CLI here * Would be better to hook the Python logging to call back into the plugin return state like here: https://pyo3.rs/v0.23.3/ecosystem/logging.html#the-python-to-rust-direction * We should only load the LineBuilder interface once in a module, rather than on every execution of a WAL plugin * More tests all around But I want to get this in so that the actual plugin and trigger system can get udated to build around this model. * refactor: PR feedback	2025-01-06 17:32:17 -05:00
Michael Gattozzi	ccda3dd3a9	feat: remove required field restriction for tables (#25738 ) This commit removes the required fields restriction when using the CLI or the API to create a new table. As users can't write via the line protocol without a field this is fine and the schema will be updated on write. This expands the test to check for the correct response code now and make sure that we can both query the empty table and write new data to it. Closes #25735	2025-01-03 18:10:56 -05:00
Jackson Newhouse	7aa8f41268	feat(processing_engine): Add CLI support for plugins and triggers (#25731 )	2025-01-03 12:11:52 -08:00
Trevor Hilton	7c7e5d77df	feat: remove table_name clause requirement for parquet files system table (#25733 ) Updated the tests to check for correct behaviour on queries to the system.parquet_files table.	2025-01-02 16:59:32 -05:00
Michael Gattozzi	636503ff46	fix: Return a 422 for limits on a partial write (#25729 ) Prior to this change we would error correctly with a 422 if a limit was hit. However, we would not send back the correct error in the case of a limit being hit that caused a partial write. This change fixes that by checking the error messages for failed lines and if one is found that caused a limit to be hit, then a 422 is returned rather than a 400 as we would have been able to process the line otherwise, but the limit was hit instead Closes #25208	2025-01-02 14:01:23 -05:00
Jackson Newhouse	29dacc318a	feat(processing_engine): Add REST API endpoints for activating and deactivating triggers. (#25711 )	2025-01-02 09:23:18 -08:00
Trevor Hilton	de227b95d9	refactor: cleanup v3 write API and series key method on catalog (#25723 ) Store the series key column names on the TableDefinitin in catalog so looking up the series key by column names is more efficient Remove the /api/v3/write API and related code/tests	2024-12-30 09:32:54 -05:00
praveen-influx	96a580f365	refactor: porting changes in pro to oss (#25712 ) - query_executor is moved to it's own module so that it's easier to swap the setup for pro. - system_tables setup is also modified	2024-12-27 15:02:22 +00:00
Michael Gattozzi	c764d37636	feat: Add json lines support to query output (#25698 )	2024-12-20 14:57:19 -05:00
Michael Gattozzi	048e45fff9	fix: change 200 to 204 for compatibility of v1/v2 (#25699 )	2024-12-20 14:56:18 -05:00
Michael Gattozzi	2a132f1edf	fix: change failed query with missing db to 404 (#25693 )	2024-12-20 11:53:13 -05:00
Trevor Hilton	d10ad87f2c	feat: write metrics (#25692 ) Added prometheus metrics to track lines written and bytes written per database. The write buffer does the tracking after validation of incoming line protocol. Tests added to verify.	2024-12-20 11:02:36 -05:00
Michael Gattozzi	e51bea65b4	feat: create DB and Tables via REST and CLI (#25687 ) * feat: create DB and Tables via REST and CLI This commit does a few things: 1. It brings the database command naming scheme for types inline with the rest of the CLI types 2. It brings the table command naming scheme for types inline with the rest of the CLI types 3. Adds tests to check that the num of dbs is not exceeded and that you cannot create more than one database with a given name. 4. Adds tests to check that you can create a table and put data into it and querying it 5. Adds tests for the CLI for both the database and table commands 6. It creates an endpoint to create databases given a JSON blob 7. It creates an endpoint to create tables given a JSON blob With this users can now create a database or table without first needing to write to the database via the line protocol! Closes #25640 Closes #25641	2024-12-19 16:01:34 -05:00
Jackson Newhouse	8bfccb74ab	feat(processing_engine): Runtime and write-back improvements (#25672 ) * Move processing engine invocation to a seperate tokio task. * Support writing back line protocol from python via insert_line_protocol(). * Update structs to work with bincode.	2024-12-17 16:38:12 -08:00
Michael Gattozzi	90b9071297	fix: Change references of Edge into OSS (#25661 ) This changes the code to reference InfluxDB 3 OSS rather than Edge which had been it's original name when we first started the project. With this we now have the code reflect what we are actually calling it. On top of this the long help text has been changed to give advice about how to actually run the code now with the bare minimum set of flags needed now as `influxdb serve` is no longer a viable command on it's own. Closes #25649	2024-12-16 11:24:35 -05:00
Jackson Newhouse	486d79d801	feat(processing_engine): initial implementation of Processing Engine plugins and triggers (#25639 )	2024-12-13 14:11:38 -08:00
Michael Gattozzi	9292a3213d	feat: Significantly decrease startup times for WAL (#25643 ) * feat: add startup time to logging output This change adds a startup time counter to the output when starting up a server. The main purpose of this is to verify whether the impact of changes actually speeds up the loading of the server. * feat: Significantly decrease startup times for WAL This commit does a few important things to speedup startup times: 1. We avoid changing an Arc<str> to a String with the series key as the From<String> impl will call with_column which will then turn it into an Arc<str> again. Instead we can just call `with_column` directly and pass in the iterator without also collecting into a Vec<String> 2. We switch to using bitcode as the serialization format for the WAL. This significantly reduces startup time as this format is faster to use instead of JSON, which was eating up massive amounts of time. Part of this change involves not using the tag feature of serde as it's currently not supported by bincode 3. We also parallelize reading and deserializing the WAL files before we then apply them in order. This reduces time waiting on IO and we eagerly evaluate each spawned task in order as much as possible. This gives us about a 189% speedup over what we were doing before. Closes #25534	2024-12-12 11:27:51 -05:00
Trevor Hilton	37219af9d4	feat: track parquet cache metrics (#25632 ) * feat: parquet cache metrics * feat: track parquet cache metrics Adds metrics to track the following in the in-memory parquet cache: * cache size in bytes (also included a fix in the calculation of that) * cache size in n files * cache hits * cache misses * cache misses while the oracle is fetching a file A test was added to check this functionality * refactor: clean up logic and fix cache removal tracking error Some logic and naming was cleaned up and the boolean to optionally track metrics on entry removal was removed, as it was incorrect in the first place: a fetching entry still has a size, which counts toward the size of the cache. So, this makes is such that anytime an entry is removed, whether its state is success or fetching, its size will be decremented from the cache size metrics. The sizing caclulations were made to be correct, and the cache metrics test was updated with more thurough assertions	2024-12-10 09:32:15 -05:00
Trevor Hilton	0bfef47ff9	refactor: move parquet cache to influxdb3_cache crate (#25630 )	2024-12-09 11:56:52 -05:00
Trevor Hilton	161c8b4eda	refactor: remove query concurrency limit (#25629 )	2024-12-09 11:06:14 -05:00
Trevor Hilton	154ff7da23	feat: LastCacheExec to track predicate pushdown in last cache queries (#25621 )	2024-12-06 10:53:19 -08:00
Trevor Hilton	9b87cd7a65	refactor: move last cache to influxdb3_cache crate (#25620 ) Moved all of the last cache implementation into the `influxdb3_cache` crate. This also splits out the implementation into three modules: - `cache.rs`: the core cache implementation - `provider.rs`: the cache provider used by the database to hold multiple caches. - `table_function.rs`: same as before, holds the DataFusion impls Tests were preserved and moved to `mod.rs`, however, they were updated to not rely on the WriteBuffer implementation, and instead use the types in the `influxdb3_cache::last_cache` module directly. This simplified the test code, while not changing any of the test assertions at all.	2024-12-05 14:04:25 -05:00
Trevor Hilton	dbb1f55b5e	chore: update core for latest sync (#25617 )	2024-12-04 14:11:13 -05:00
Michael Gattozzi	d2fbd65a44	feat: Deny extra tags on write APIs (#25596 ) This commit does three important major changes: 1. We will deny writes to the v1, v2, and v3 write apis that add new tags in subsequent writes after the first write 2. We make every table have a series key by default now 3. We enfore sorting order by the series key which is the order the keys came in With these changes we have consistentcy across the various write apis and can make optimizations and future features with the assumption we have a series key. Closes #25585	2024-12-03 12:10:26 -05:00
praveen-influx	43755c2d9c	feat: sys events store added (#25603 ) This commit introduces basic store for sys events and the backing ring buffer. Since the buffer needs to hold arbitrary data, it uses `Box<dyn Any>` closes: https://github.com/influxdata/influxdb/issues/25581	2024-12-02 10:55:37 +00:00
Trevor Hilton	a01fd16d4d	fix: plan queries on DF threadpool to not block IO in REST API (#25604 )	2024-11-29 16:03:17 -05:00
Trevor Hilton	9ead1dfe4b	feat: meta_caches system table (#25593 ) This adds a new system table "meta_caches" that allows users to view the state of their metadata caches on a per-db basis An integration test was added to verify that it works.	2024-11-28 08:57:02 -05:00
Trevor Hilton	234d37329a	feat: metacache REST APIs to create and delete (#25587 )	2024-11-27 08:41:46 -05:00
praveen-influx	bfa0e71558	feat: make query executor as trait object (#25591 ) * feat: make query executor as trait object This commit moves `QueryExecutorImpl` behind a `dyn` (trait object) as we have other impls in core for `QueryExecutor` and this will keep both pro and OSS traits in sync * chore: fix cargo audit failures - address https://rustsec.org/advisories/RUSTSEC-2024-0399.html by running `cargo update --precise 0.23.18 --package rustls@0.23.14` - address yanked version of `url` crate (2.5.3) by running `cargo update -p url`	2024-11-26 17:18:22 +00:00
praveen-influx	3cde24feb4	feat: delete table (#25572 ) This commit allows deleting (soft) a table. For an user, following command will allow soft deleting a table (bar) in db (foo) ``` influxdb3 table delete --dbname foo --table bar --host $host ``` - Added `soft_delete_table` to `DatabaseManager` trait, which already hosts `soft_delete_database` method. The code roughly follows the same flow as db delete. Although like db schema, it does clone on write because the reference is behind an Arc, `Arc::make_mut` is used in this change. - Moved db delete related cli parser under "manage" module that has both db and table delete functionality - Some minor tidyups (removing unused methods, renaming method so that the order in name matches actual return type eg. `table_id_and_schema`, should return (id, schema) and not (schema, id)) closes: https://github.com/influxdata/influxdb/issues/25561	2024-11-22 08:42:45 +00:00
praveen-influx	33c2d47ba9	feat: drop/delete database (#25549 ) * feat: drop/delete database This commit allows soft deletion of database using `influxdb3 database delete <db_name>` command. The write buffer and last value cache are cleared as well. closes: https://github.com/influxdata/influxdb/issues/25523 * feat: reuse same code path when deleting database - In previous commit, the deletion of database immediately triggered clearing last cache and query buffer. But on restarts same logic had to be repeated to allow deleting database when starting up. This commit removes immediate deletion by explicitly calling necessary methods and moves the logic to `apply_catalog_batch` which already applies `CatalogOp` and also clearing cache and buffer in `buffer_ops` method which has hooks to call other places. closes: https://github.com/influxdata/influxdb/issues/25523 * feat: use reqwest query api for query param Co-authored-by: Trevor Hilton <thilton@influxdata.com> * feat: include deleted flag in DatabaseSnapshot - `DatabaseSchema` serialization/deserialization is delegated to `DatabaseSnapshot`, so the `deleted` flag should be included in `DatabaseSnapshot` as well. - insta test snapshots fixed closes: https://github.com/influxdata/influxdb/issues/25523 * feat: address PR comments + tidy ups --------- Co-authored-by: Trevor Hilton <thilton@influxdata.com>	2024-11-19 16:08:14 +00:00
praveen-influx	814eb31309	chore: update core deps (#25532 ) * chore: update core deps - arrow/parquet deps are patched (as in core) - three specific code changes to cope with changes in core crates - TransitionPartitionId, use `from_parts` instead of `new` - arrow buffers can take &[u8] directly without `to_vec()`/`vec!` (used only in tests) - `schema` and `influxdb_line_protocol` crates need `v3` feature enabled * chore: update deny.toml * chore: formatting and deny toml changes Unicode-3.0 license is added to allowed licenses list, without it end up with 19 errors (`zerovec`, `zerovec-derive` etc.) * chore: address PR feedback - move enabling v3 feature to root Cargo.toml - added the upstream PR for datafusion-common that introduced RUSTSEC-2024-0384	2024-11-12 16:07:31 +00:00
praveen-influx	c2b8a3a355	feat: add column names to last cache sys table (#25521 ) * feat: add column names to last cache sys table closes: https://github.com/influxdata/influxdb/issues/25511 * feat: move all `get_by_id` methods to take reference in schema	2024-11-08 16:08:30 +00:00
Trevor Hilton	391b67f9ab	feat: configurable last cache eviction (#25520 ) * refactor: make last cache eviction optional This changes how the last cache is evicted. It will no longer run eviction on writes to the cache, instead, there is an optional method to create a last cache provider that will run eviction in a background task on a specified interval. Otherwise, when records are produced from the cache, only those that have not expired will be produced. This should reduce locks on the cache and hopefully improve performance. * feat: configurable last cache eviction interval * docs: clean up var names, code docs, and comments	2024-11-06 09:59:17 -05:00
Trevor Hilton	d26a73802a	refactor: move to `ColumnId` and `Arc<str>` as much as possible (#25495 ) Closes #25461 _Note: the first three commits on this PR are from https://github.com/influxdata/influxdb/pull/25492_ This PR makes the switch from using names for columns to the use of `ColumnId`s. Where column names are used, they are represented as `Arc<str>`. This impacts most components of the system, and the result is a fairly sizeable change set. The area where the most refactoring was needed was in the last-n-value cache. One of the themes of this PR is to rely less on the arrow `Schema` for handling the column-level information, and tracking that info in our own `ColumnDefinition` type, which captures the `ColumnId`. I will summarize the various changes in the PR below, and also leave some comments in-line in the PR. ## Switch to `u32` for `ColumnId` The `ColumnId` now follows the `DbId` and `TableId`, and uses a globally unique `u32` to identify all columns in the database. This was a change from using a `u16` that was only unique within the column's table. This makes it easier to follow the patterns used for creating the other identifier types when dealing with columns, and should reduce the burden of having to manage the state of a table-scoped identifier. ## Changes in the WAL/Catalog * `WriteBatch` now contains no names for tables or columns and purely uses IDs * This PR relies on `IndexMap` for `_Id`-keyed maps so that the order of elements in the map is consistent. This has important implications, namely, that when iterating over an ID map, the elements therein will always be produced in the same order which allows us to make assertions on column order in a lot of our tests, and allows for the re-introduction of `insta` snapshots for serialization tests. This map type provides O(1) lookups, but also provides _fast_ iteration, which should help when serializing these maps in write batches to the WAL. * Removed the need to serialize the bi-directional maps for `DatabaseSchema`/`TableDefinition` via use of `SerdeVecMap` (see comments in-line) * The `tables` map in `DatabaseSchema` no stores an `Arc<TableDefinition>` so that the table definition can be shared around more easily. This meant that changes to tables in the catalog need to do a clone, but we were already having to do a clone for changes to the DB schema. * Removal of the `TableSchema` type and consolidation of its parts/functions directly onto `TableDefinition` * Added the `ColumnDefinition` type, which represents all we need to know about a column, and is used in place of the Arrow `Schema` for column-level meta-info. We were previously relying heavily on the `Schema` for iterating over columns, accessing data types, etc., but this gives us an API that we have more control over for our needs. The `Schema` is still held at the `TableDefinition` level, as it is needed for the query path, and is maintained to be consistent with what is contained in the `ColumnDefinition`s for a table. ## Changes in the Last-N-Value Cache * There is a bigger distinction between caches that have an explicit set of value columns, and those that accept new fields. The former should be more performant. * The Arrow `Schema` is managed differently now: it used to be updated more than it needed to be, and now is only updated when a row with new fields is pushed to a cache that accepts new fields. ## Changes in the write-path * When ingesting, during validation, field names are qualified to their associated column ID	2024-11-01 16:42:57 -04:00
praveen-influx	cdb1e862b9	feat: set tcp_nodelay in hyper builder (#25488 ) closes: https://github.com/influxdata/influxdb/issues/25466	2024-10-24 21:48:24 +01:00
Trevor Hilton	ce9276d96d	refactor: changes needed for IDs in pro (#25479 ) * refactor: roll back addition of DatabaseSchemaProvider trait * refactor: make parquet metrics optional in telemetry for pro * refactor: make ParquetFileId Hash * refactor: test harness logging	2024-10-21 15:17:02 -04:00
Trevor Hilton	10b6a2810d	refactor: separte catalog schema API (#25468 ) Separate out methods of the Catalog API that are used on the query side into a new trait `DatabaseSchemaProvider`. The new trait includes methods from the Catalog that get the underlying `DatabaseSchema` or interact with names/IDs. This will allow for a separate implementation of the Catalog for pro that only needs to hold a replicated/combined view in-memory of one or more catalogs without the need to do persistence that a write buffer's catalog needs to do. While in there I also switched the `QueryExecutorImpl::new` method to take an args struct to avoid the clippy lint.	2024-10-16 11:42:07 -04:00
Trevor Hilton	ead6d27c5f	refactor: improvements to the catlog API (#25463 )	2024-10-11 13:44:43 -04:00
Jamie Strandboge	0835093c78	feat(circleci): add inclusivity checks (#25437 ) * feat(circleci): add inclusivity checks * chore(circleci): adjust package-validation for inclusive language * chore: update tests for inclusive language	2024-10-09 08:01:31 -05:00
praveen-influx	1f1125c767	refactor: udpate docs and tests for the telemetry crate (#25432 ) - Introduced traits, `ParquetMetrics` and `SystemInfoProvider` to enable writing easier tests - Uses mockito for code that depends on reqwest::Client and also uses mockall to generally mock any traits like `SystemInfoProvider` - Minor updates to docs	2024-10-08 15:45:13 +01:00
Michael Gattozzi	c4534b06da	feat: move Table Id/Name mapping into DB Schema (#25436 )	2024-10-08 10:03:55 -04:00
Michael Gattozzi	eeb1aa7905	feat: swap over to DbId and TableId everywhere (#25421 ) * feat: Add TableId and ColumnId * feat: swap over to DbId and TableId everywhere This commit swaps us over to using the DbId and TableId types everywhere for our internal systems. Anywhere that's external facing, such as names for last cache tables or line protocol parsing, use names. In these cases we have the `Catalog` which keeps a map of TableIds and DbIds in a bidirectional mapping for easy lookup i.e. id <-> names. While in essence the change itself isn't that complicated given the nature of how much we depended on names for things, the changes end up being quite invasive and extensive. Luckily it shouldn't be too hard to review. Note this does not add the column ids which will be done in a follow up PR. Closes #25375 Closes #25403 Closes #25404 Closes #25405 Closes #25412 Closes #25413	2024-10-03 14:47:46 -04:00
praveen-influx	8ccb580162	feat: telemetry report for parquet metrics (#25425 ) - added mechanism within PersistedFile to expose parquet file related metrics. The details are updated when new snapshot is generated and also when all snapshots are loaded when the process starts up - at the point of creating the telemetry payload these parquet metrics are looked up before sending it to the server. Closes: https://github.com/influxdata/influxdb/issues/25418	2024-10-03 15:11:40 +01:00
praveen-influx	72dcd1866f	feat(telemetry): adds reads and writes (#25409 ) - instrumented code to get read and write measurement - introduced EventsBucket for collection of reads/writes - sampler now samples every minute for all metrics (including reads/writes) - other tidy ups closes: https://github.com/influxdata/influxdb/issues/25372	2024-10-01 18:34:00 +01:00
Trevor Hilton	a05c3fe87b	test: check parquet cache in the write buffer (#25411 ) * test: check parquet cache in the write buffer Checked that the parquet cache will serve queries when chunks are requested from the write buffer. The added test also checks for get_range requests made to the object store, which are typically made by DataFusion to infer schema for parquet files. * refactor: make parquet cache optional on write buffer * test: add test to verify parquet cache function This makes the parquet cache optional at the write buffer level, and adds a test that verifies that the cache catches and prevents requests to the object store in the event of a cache hit.	2024-09-30 14:54:12 -04:00
Trevor Hilton	4184a331ea	refactor: parquet cache with less locking (#25389 ) Closes #25382 Closes #25383 This refactors the parquet cache to use less locking by switching from using the `clru` crate to a hand-rolled cache implementation. The new cache still acts as an LRU, but it uses atomics to track hit-time per entry, and handles pruning in a separate process that is decoupled from insertion/gets to the cache. The `Cache` type uses a [`DashMap`](https://docs.rs/dashmap/latest/dashmap/struct.DashMap.html) internally to store cache entries. This should help reduce lock contention, and also has the added benefit of not requiring mutability to insert into _or_ get from the map. The cache maps an `object_store::Path` to a `CacheEntry`. On a hit, an entry will have its `hit_time` (an `AtomicI64`) incremented. During a prune operation, entries that have the oldest hit times will be removed from the cache. See the `Cache::prune` method for details. The cache is setup with a memory _capacity_ and a _prune percent_. The cache tracks memory used when entries are added, based on their _size_, and when a prune is invoked in the background, if the cache has exceeded its capacity, it will prune `prune_percent * cache.len()` entries from the cache. Two tests were added: * `cache_evicts_lru_when_full` to check LRU behaviour of the cache * `cache_hit_while_fetching` to check that a cache entry hit while a request is in flight to fetch that entry will not result in extra calls to the underlying object store	2024-09-27 11:59:17 -04:00
praveen-influx	70643d0136	feat(auth): allow Token or Bearer as valid schemes (#25397 ) closes: https://github.com/influxdata/influxdb/issues/25394	2024-09-27 13:40:28 +01:00
Trevor Hilton	9c71b3ce25	feat: memory-cached object store for parquet files (#25377 ) Part of #25347 This sets up a new implementation of an in-memory parquet file cache in the `influxdb3_write` crate in the `parquet_cache.rs` module. This module introduces the following types: * `MemCachedObjectStore` - a wrapper around an `Arc<dyn ObjectStore>` that can serve GET-style requests to the store from an in-memory cache * `ParquetCacheOracle` - an interface (trait) that can accept requests to create new cache entries in the cache used by the `MemCachedObjectStore` * `MemCacheOracle` - implementation of the `ParquetCacheOracle` trait ## `MemCachedObjectStore` This takes inspiration from the [`MemCacheObjectStore` type](`1eaa4ed5ea/object_store_mem_cache/src/store.rs (L205-L213)`) in core, but has some different semantics around its implementation of the `ObjectStore` trait, and uses a different cache implementation. The reason for wrapping the object store is that this ensures that any GET-style request being made for a given object is served by the cache, e.g., metadata requests made by DataFusion. The internal cache comes from the [`clru` crate](https://crates.io/crates/clru), which provides a least-recently used (LRU) cache implementation that allows for weighted entries. The cache is initialized with a capacity and entries are given a weight on insert to the cache that represents how much of the allotted capacity they will take up. If there isn't enough room for a new entry on insert, then the LRU item will be removed. ### Limitations of `clru` The `clru` crate conveniently gives us an LRU eviction policy but its API may put some limitations on the system: * gets to the cache require an `&mut` reference, which means that the cache needs to be behind a `Mutex`. If this slows down requests through the object store, then we may need to explore alternatives. * we may want more sophisticated eviction policies than a straight LRU, i.e., to favour certain tables over others, or files that represent recent data over those that represent old data. ## `ParquetCacheOracle` / `MemCacheOracle` The cache oracle is responsible for handling cache requests, i.e., to fetch an item and store it in the cache. In this PR, the oracle runs a background task to handle these requests. I defined this as a trait/struct pair since the implementation may look different in Pro vs. OSS.	2024-09-24 10:58:15 -04:00
Michael Gattozzi	54d209d0bf	feat: Add u32 ID for Databases (#25302 ) * feat: Remove lock for FileId tests Since we now are using cargo-nextest in CI we can remove the locks used in the FileId tests to make sure that we have no race conditions * feat: Add u32 ID for Databases This commit adds a new DbId for databases. It also updates paths to use that id as part of the name. When starting up the WriteBuffer we apply the DbId from the persisted snapshot much like we do for ParquetFileId's This introduces the influxdb3_id crate to avoid circular deps with ids. The ParquetFileId should also be moved into this crate, but it's outside the scope of this change. Closes #25301	2024-09-18 11:44:04 -04:00
praveen-influx	245b49ae0e	feat: add instance id and host id to catalog (#25343 ) - uses Arc<str> to represent create once and read everywhere type of string - updated snapshots for insta asserts, uses redaction to hardcode randomly generated UUID strings - added methods to catalog to expose instace and host ids Closes: https://github.com/influxdata/influxdb/issues/25315	2024-09-17 16:32:21 +01:00
Paul Dix	f8b6cfac5b	refactor: Rename level 0 to gen1 to match compaction wording (#25317 )	2024-09-12 15:57:30 -04:00
Trevor Hilton	ad2ca83d72	chore: sync to latest core (#25284 )	2024-09-06 13:49:38 -04:00
Trevor Hilton	cd23be6e5c	test: repro for dropped wal files during snapshot (#25276 ) * test: repro for dropped wal files during snapshot This commit provides a reproducer for an issue in the snapshotting process whereby WAL files are removed for writes that have not been persisted yet. * fix: do not snapshot most recent WAL period This addresses #25277 Snapshots that are triggered when the number of WAL periods in the tracker grows to be >= 3x the snapshot size will not include the most recent wal period, and doing so was removing WAL files containing data that was not yet persisted. * docs: add doc comment to reproducer test * fix: broken parquet_files system table test * fix: broken snapshot_tracker test * fix: broken write_buffer test * refactor: remove redundant helper function * test: add another snapshot test to write_buffer * test: future writes do not get dropped on restart	2024-09-03 20:21:33 -04:00
Trevor Hilton	3f7b0e655c	refactor: move catalog and last cache initialization out of write buffer (#25273 ) * refactor: add catalog as dep to influxdb3 * refactor: move catalog and last cache initialization out of write buffer The Write buffer used to handle initialization of the catalog and last n value cache. This commit moves that logic out, so that both can be initialized independently, and injected into the write buffer. This is to enable downstream changes that will need to make sharing the catalog and last cache possible.	2024-08-27 16:41:40 -04:00
Trevor Hilton	e0e0075766	refactor: use more `dyn Trait` in write buffer (#25264 ) * refactor: use dyn traits in WriteBufferImpl This changes the WriteBufferImpl to use a dyn TimeProvider instead of a generic in its type signature. The Server type now uses a dyn WriteBuffer instead of using a generic in its type signature, and the ServerBuilder was updated to accommodate this accordingly. These chages were to make downstream code changes more seamless. * refactor: make some items pub This makes functions on the QueryableBuffer and LastCache pub so that they can be used downstream.	2024-08-23 14:21:20 -04:00
Trevor Hilton	cbb7bc5901	refactor: remove Persister trait in favour of concrete impl (#25260 ) The Persister trait was only implemented by a single type, because the underlying ObjectStore interface has several ways of being mocked, we mock that instead of the Persister interface. This commit removes the Persister trait, and moves its interface/impl directly on a single Persister type in the persister module of the influxdb3_write crate. deny.toml had some incorrect field names in license.exceptions, those were fixed from 'crate' to 'name'.	2024-08-22 10:41:33 -04:00
Paul Dix	8bcc7522d0	feat: Add last cache create/delete to WAL (#25233 ) * feat: Add last cache create/delete to WAL This moves the LastCacheDefinition into the WAL so that it can be serialized there. This ended up being a pretty large refactor to get the last cache creation to work through the WAL. I think I also stumbled on a bug where the last cache wasn't getting initialized from the catalog on reboot so that it wouldn't actually end up caching values. The refactored last cache persistence test in write_buffer/mod.rs surfaced this. Finally, I also had to update the WAL so that it would persist if there were only catalog updates and no writes. Fixes #25203 * fix: typos	2024-08-09 05:46:35 -07:00
Trevor Hilton	2082135077	fix: main (#25231 )	2024-08-08 11:56:09 -04:00
Paul Dix	05ab730ae6	refactor: Make Level0Duration part of WAL (#25228 ) * refactor: Make Level0Duration part of WAL I noticed this during some testing and cleanup with other PRs. The WAL had its own level_0_duration and the write buffer had a different one, which would cause some weird problems if they weren't the same. This refactors Level0Duration to be in the WAL and fixes up the tests. As an added bonus, this surfaced a bug where multiple L0 blocks getting persisted in the same snapshot wasn't supported. So now snapshot details can have many files per table. * fix: have persisted files always return in descending data time order * fix: sort record batches for test verification	2024-08-08 09:47:21 -04:00
Trevor Hilton	4067c91be0	fix: un-pub QueryableBuffer and fix compile errors (#25230 )	2024-08-08 09:39:12 -04:00
Trevor Hilton	7474c0b3b4	feat: add `system.parquet_files` table (#25225 ) This extends the system tables available with a new `parquet_files` table which will list the parquet files associated with a given table in a database. Queries to system.parquet_files must provide a table_name predicate to specify the table name of interest. The files are accessed through the QueryableBuffer. In addition, a test was added to check success and failure modes of the new system table query. Finally, the Persister trait had its associated error type removed. This was somewhat of a consequence of how I initially implemented this change, but I felt cleaned the code up a bit, so I kept it in the commit.	2024-08-08 08:46:26 -04:00
Trevor Hilton	b0beab5b0c	feat: use host identifier prefix in object store paths (#25224 ) This enforces the use of a host identifier prefix in all object store paths (currently, for parquet files, catalog files, and snapshot files). The persister retains the host identifier prefix, and uses it when constructing paths. The WalObjectStore also holds the host identifier prefix, so that it can use it when saving and loading WAL files. The influxdb3 binary requires a new argument 'host-id' to be passed that is used to specify the prefix.	2024-08-07 16:23:36 -04:00
Paul Dix	2b8fc7b44e	refactor: Move Catalog into influxdb3_catalog crate (#25210 ) * refactor: Move Catalog into influxdb3_catalog crate This moves the catalog and its serialization logic into its own crate. This is a precursor to recording more catalog modifications into the WAL. Fixes #25204 * fix: cargo update * fix: add version = 2 to deny.toml * fix: update deny.toml * fix: add CCO to deny.toml	2024-08-02 16:04:12 -04:00
Paul Dix	3265960010	refactor: implement new wal and refactor write buffer (#25196 ) * feat: refactor WAL and WriteBuffer There is a ton going on here, but here are the high level things. This implements a new WAL, which is backed entirely by object store. It then updates the WriteBuffer to be able to work with how the new WAL works, which also required an update to how the Catalog is modified and persisted. The concept of Segments has been removed. Previously there was a separate WAL per segment of time. Instead, there is now a single WAL that all writes and updates flow into. Data within the write buffer is organized by Chunk(s) within tables, which is based on the timestamp of the row data. These are known as the Level0 files, which will be persisted as Parquet into object store. The default chunk duration for level 0 files is 10 minutes. The WAL is written as single files that get created at the configured WAL flush interval (1s by default). After a certain number of files have been created, the server will attempt to snapshot the WAL (default is to snapshot the first 600 files of the WAL after we have 900 total, i.e. snapshot 10 minutes of WAL data). The design goal with this is to persist 10 minute chunks of data that are no longer receiving writes, while clearing out old WAL files. This works if data getting written in around "now" with no more than 5 minutes of delay. If we continue to have delayed writes, a snapshot of all data will be forced in order to clear out the WAL and free up memory in the buffer. Overall, this structure of a single wal, with flushes and snapshots and chunks in the queryable buffer led to a simpler setup for the write buffer overall. I was able to clear out quite a bit of code related to the old segment organization. Fixes #25142 and fixes #25173 * refactor: address PR feedback * refactor: wal to replay and background flush on new * chore: remove stray println	2024-08-01 15:04:15 -04:00

1 2 3 4 5

205 Commits (fix/docker-python-deps-earlier)