influxdb

Commit Graph

Author	SHA1	Message	Date
Paul Dix	4bfd95d068	feat: print plugin logs to the server (#25812 )	2025-01-12 12:33:09 -05:00
Paul Dix	491a37b0d4	feat: Update create plugin to use server file (#25803 ) This updates the create plugin API and CLI so that it doesn't take the pugin code, but instead takes a file name of a file that must be in the plugin-dir of the server. It returns an error if the plugin-dir is not configured or if the file isn't there. Also updates the WAL and catalog so that it doesn't store the plugin code directly. The code is read from disk one time when the plugin runs. Closes #25797	2025-01-11 21:02:51 -05:00
Trevor Hilton	0bdc2fa953	chore: patch enterprise back to core (#25798 )	2025-01-11 17:26:41 -05:00
Paul Dix	e8422a240a	feat: Wire up arguments to wal plugin trigger (#25783 ) This allows the user to specify arguments that will be passed to each execution of a wal plugin trigger. The CLI test was updated to check this end to end. Closes #25655	2025-01-10 16:58:18 -05:00
Paul Dix	7230148b58	feat: Update WAL plugin for new structure (#25777 ) * feat: Update WAL plugin for new structure This ended up being a very large change set. In order to get around circular dependencies, the processing engine had to be moved into its own crate, which I think is ultimately much cleaner. Unfortunately, this required changing a ton of things. There's more testing and things to add on to this, but I think it's important to get this through and build on it. Importantly, the processing engine no longer resides inside the write buffer. Instead, it is attached to the HTTP server. It is now able to take a query executor, write buffer, and WAL so that the full range of functionality of the server can be exposed to the plugin API. There are a bunch of system-py feature flags littered everywhere, which I'm hoping we can remove soon. * refactor: PR feedback	2025-01-10 05:52:33 -05:00
Paul Dix	2d18a61949	feat: Add query API to Python plugins (#25766 ) This ended up being a couple things rolled into one. In order to add a query API to the Python plugin, I had to pull the QueryExecutor trait out of server into a place so that the python crate could use it. This implements the query API, but also fixes up the WAL plugin test CLI a bit. I've added a test in the CLI section so that it shows end-to-end operation of the WAL plugin test API and exercise of the entire Plugin API. Closes #25757	2025-01-09 20:13:20 -05:00
praveen-influx	6e2e39cd4c	feat: snapshot when wal buffer is empty (#25765 ) * feat: snapshot when wal buffer is empty - This commit changes the functionality to allow snapshots to happen even when wal buffer is empty. For snapshots wal periods are still required but not the wal buffer. To allow this, we write a no-op into wal file with snapshot details. This enables force snapshotting functionality closes: https://github.com/influxdata/influxdb/issues/25685 * refactor: address PR feedback	2025-01-09 12:12:37 +00:00
Trevor Hilton	dfc853d903	feat: handle params in request body for /query API (#25762 ) Closes #25749 This changes the `/query` API handler so that the parameters can be passed in either the request URI or in the request body for either a `GET` or `POST` request. Parameters can be specified in the URI, the body, or both; if they are specified in both places, those in the body will take precedent. Error variants in the HTTP server code related to missing request parameters were updated to return `400` status.	2025-01-07 19:53:29 -05:00
Michael Gattozzi	f793d31f63	feat: Cleanup CLI flags for InfluxDB 3 Core (#25737 ) This makes quite a few major changes to our CLI and how users interact with it: 1. All commands are now in the form <verb> <noun> this was to make the commands consistent. We had last-cache as a noun, but serve as a verb in the top level. Given that we could only create or delete All noun based commands have been move under a create and delete command 2. --host short form is now -H not -h which is reassigned to -h/--help for shorter help text and is in line with what users would expect for a CLI 3. Only the needed items from clap_blocks have been moved into `influxdb3_clap_blocks` and any IOx specific references were changed to InfluxDB 3 specific ones 4. References to InfluxDB 3.0 OSS have been changed to InfluxDB 3 Core in our CLI tools 5. --dbname has been changed to --database to be consistent with --table in many commands. The short -d flag still remains. In the create/ delete command for the database however the name of the database is a positional arg e.g. `influxbd3 create database foo` rather than `influxdb3 database create --dbname foo` 6. --table has been removed from the delete/create command for tables and is now a positional arg much like database 7. clap_blocks was removed as dependency to avoid having IOx specific env vars 8. --cache-name is now an optional positional arg for last_cache and meta_cache 9. last-cache/meta-cache commands are now last_cache and meta_cache respectively Unfortunately we have quite a few options to run the software and I couldn't cut down on them, but at least with this commands and options will be more discoverable and we have full control over our CLI options now. Closes #25646	2025-01-06 18:51:55 -05:00
Paul Dix	1ce6a24c3f	feat: Implement WAL plugin test API (#25704 ) * feat: Implement WAL plugin test API This implements the WAL plugin test API. It also introduces a new API for the Python plugins to be called, get their data, and call back into the database server. There are some things that I'll want to address in follow on work: * CLI tests, but will wait on #25737 to land for a refactor of the CLI here * Would be better to hook the Python logging to call back into the plugin return state like here: https://pyo3.rs/v0.23.3/ecosystem/logging.html#the-python-to-rust-direction * We should only load the LineBuilder interface once in a module, rather than on every execution of a WAL plugin * More tests all around But I want to get this in so that the actual plugin and trigger system can get udated to build around this model. * refactor: PR feedback	2025-01-06 17:32:17 -05:00
Michael Gattozzi	d78756f02f	feat: allow n/u for precision in v1/v2 write apis (#25742 ) Prior to this change we would deny writes that used n and u for the precision argument when doing writes. We only accepted ns and us for those apis. However, to be backwards compatible we would need to enable accepting writes with n and u. This is mostly just upgrading our deps as this was a change that landed in IOx first. We test that it works for our code by adding test cases for their precision in this repo.	2025-01-03 19:23:33 -05:00
Trevor Hilton	03ea565802	feat: cli arg to specify max parquet fanout (#25714 ) This allows the `max_parquet_fanout` to be specified in the CLI for the `influxdb3 serve` command. This could be done previously via the `--datafusion-config` CLI argument, but the drawbacks to that were: 1. that is a fairly advanced option given the available key/value pairs are not well documented 2. if `iox.max_parquet_fanout` was not provided to that argument, the default would be set to `40` This PR maintains the existing `--datafusion-config` CLI argument (with one caveat, see below) which allows users to provide a set key/value pairs that will be used to build the internal DataFusion config, but in addition provides the `--datafusion-max-parquet-fanout` argument: ``` --datafusion-max-parquet-fanout <MAX_PARQUET_FANOUT> When multiple parquet files are required in a sorted way (e.g. for de-duplication), we have two options: 1. In-mem sorting: Put them into `datafusion.target_partitions` DataFusion partitions. This limits the fan-out, but requires that we potentially chain multiple parquet files into a single DataFusion partition. Since chaining sorted data does NOT automatically result in sorted data (e.g. AB-AB is not sorted), we need to preform an in-memory sort using `SortExec` afterwards. This is expensive. 2. Fan-out: Instead of chaining files within DataFusion partitions, we can accept a fan-out beyond `target_partitions`. This prevents in-memory sorting but may result in OOMs (out-of-memory) if the fan-out is too large. We try to pick option 2 up to a certain number of files, which is configured by this setting. [env: INFLUXDB3_DATAFUSION_MAX_PARQUET_FANOUT=] [default: 1000] ``` with the default value of `1000`, which will override the core `iox_query` default of `40`. A test was added to check that this is propagated down to the `IOxSessionContext` that is used during queries. The only change to the `datafusion-config` CLI argument was to rename `INFLUXDB_IOX` in the environment variable to `INFLUXDB3`: ``` --datafusion-config <DATAFUSION_CONFIG> Provide custom configuration to DataFusion as a comma-separated list of key:value pairs. # Example ```text --datafusion-config "datafusion.key1:value1, datafusion.key2:value2" ``` [env: INFLUXDB3_DATAFUSION_CONFIG=] [default: ] ```	2024-12-27 12:42:30 -05:00
praveen-influx	3f678678d7	chore: update core dependencies (#25708 ) - one notable change is `make_object_store` from clap_blocks has been removed. Instead use `ObjectStoreConfig::make_object_store()`	2024-12-24 14:21:59 +00:00
Jackson Newhouse	8bfccb74ab	feat(processing_engine): Runtime and write-back improvements (#25672 ) * Move processing engine invocation to a seperate tokio task. * Support writing back line protocol from python via insert_line_protocol(). * Update structs to work with bincode.	2024-12-17 16:38:12 -08:00
Paul Dix	31b9209dd6	fix: Snapshot QueryableBuffer error (#25673 ) Fixes bug in queryable buffer where if a block of data was missing one of the columns defined in a table sort key, the creation of the logical plan to sort and dedupe the data would fail, causing a panic. Fixes #25670	2024-12-17 16:57:07 -05:00
Trevor Hilton	7d92b75731	feat: add influxdb3_clap_blocks crate with runtime config (#25665 ) * feat: add influxdb3_clap_blocks with runtime config Added a new workspace crate `influxdb3_clap_blocks` which will be a starting point for adding InfluxDB 3 OSS/Pro specific CLI configuration that no longer references IOx, and allows for us to trim out unneeded configurations for the monolithic InfluxDB 3. Other than changing references from IOX to INFLUXDB3, this makes one important change: it enables IO on the DataFusion runtime. This, for now, is an experimental change to see if we can relieve some concurrency issues that we have been experiencing. * chore: add observability deps for windows	2024-12-16 15:31:55 -05:00
Jackson Newhouse	486d79d801	feat(processing_engine): initial implementation of Processing Engine plugins and triggers (#25639 )	2024-12-13 14:11:38 -08:00
Michael Gattozzi	9292a3213d	feat: Significantly decrease startup times for WAL (#25643 ) * feat: add startup time to logging output This change adds a startup time counter to the output when starting up a server. The main purpose of this is to verify whether the impact of changes actually speeds up the loading of the server. * feat: Significantly decrease startup times for WAL This commit does a few important things to speedup startup times: 1. We avoid changing an Arc<str> to a String with the series key as the From<String> impl will call with_column which will then turn it into an Arc<str> again. Instead we can just call `with_column` directly and pass in the iterator without also collecting into a Vec<String> 2. We switch to using bitcode as the serialization format for the WAL. This significantly reduces startup time as this format is faster to use instead of JSON, which was eating up massive amounts of time. Part of this change involves not using the tag feature of serde as it's currently not supported by bincode 3. We also parallelize reading and deserializing the WAL files before we then apply them in order. This reduces time waiting on IO and we eagerly evaluate each spawned task in order as much as possible. This gives us about a 189% speedup over what we were doing before. Closes #25534	2024-12-12 11:27:51 -05:00
Trevor Hilton	37219af9d4	feat: track parquet cache metrics (#25632 ) * feat: parquet cache metrics * feat: track parquet cache metrics Adds metrics to track the following in the in-memory parquet cache: * cache size in bytes (also included a fix in the calculation of that) * cache size in n files * cache hits * cache misses * cache misses while the oracle is fetching a file A test was added to check this functionality * refactor: clean up logic and fix cache removal tracking error Some logic and naming was cleaned up and the boolean to optionally track metrics on entry removal was removed, as it was incorrect in the first place: a fetching entry still has a size, which counts toward the size of the cache. So, this makes is such that anytime an entry is removed, whether its state is success or fetching, its size will be decremented from the cache size metrics. The sizing caclulations were made to be correct, and the cache metrics test was updated with more thurough assertions	2024-12-10 09:32:15 -05:00
Trevor Hilton	0bfef47ff9	refactor: move parquet cache to influxdb3_cache crate (#25630 )	2024-12-09 11:56:52 -05:00
Trevor Hilton	ef3599d7ce	test: metadata cache query using JSON format (#25626 )	2024-12-06 15:36:49 -05:00
Trevor Hilton	9b87cd7a65	refactor: move last cache to influxdb3_cache crate (#25620 ) Moved all of the last cache implementation into the `influxdb3_cache` crate. This also splits out the implementation into three modules: - `cache.rs`: the core cache implementation - `provider.rs`: the cache provider used by the database to hold multiple caches. - `table_function.rs`: same as before, holds the DataFusion impls Tests were preserved and moved to `mod.rs`, however, they were updated to not rely on the WriteBuffer implementation, and instead use the types in the `influxdb3_cache::last_cache` module directly. This simplified the test code, while not changing any of the test assertions at all.	2024-12-05 14:04:25 -05:00
praveen-influx	7211e8a96c	feat: move ring buffer to use array instead of vec (#25616 ) In this commit the vec backing the buffer is swapped for an array. Criterion benchmarks were added to compare the perf to make sure it has not made it worse. The vec implementation has been removed after the benchmarks done locally	2024-12-04 21:05:07 +00:00
Trevor Hilton	dbb1f55b5e	chore: update core for latest sync (#25617 )	2024-12-04 14:11:13 -05:00
praveen-influx	43755c2d9c	feat: sys events store added (#25603 ) This commit introduces basic store for sys events and the backing ring buffer. Since the buffer needs to hold arbitrary data, it uses `Box<dyn Any>` closes: https://github.com/influxdata/influxdb/issues/25581	2024-12-02 10:55:37 +00:00
Trevor Hilton	234d37329a	feat: metacache REST APIs to create and delete (#25587 )	2024-11-27 08:41:46 -05:00
praveen-influx	bfa0e71558	feat: make query executor as trait object (#25591 ) * feat: make query executor as trait object This commit moves `QueryExecutorImpl` behind a `dyn` (trait object) as we have other impls in core for `QueryExecutor` and this will keep both pro and OSS traits in sync * chore: fix cargo audit failures - address https://rustsec.org/advisories/RUSTSEC-2024-0399.html by running `cargo update --precise 0.23.18 --package rustls@0.23.14` - address yanked version of `url` crate (2.5.3) by running `cargo update -p url`	2024-11-26 17:18:22 +00:00
Trevor Hilton	8e23032ceb	feat: add metadata cache provider with APIs for write and query (#25566 ) This adds the MetaDataCacheProvider for managing metadata caches in the influxdb3 instance. This includes APIs to create caches through the WAL as well as from a catalog on initialization, to write data into the managed caches, and to query data out of them. The query side is fairly involved, relying on Datafusion's TableFunctionImpl and TableProvider traits to make querying the cache using a user-defined table function (UDTF) possible. The predicate code was modified to only support two kinds of predicates: IN and NOT IN, which simplifies the code, and maps nicely with the DataFusion LiteralGuarantee which we leverage to derive the predicates from the incoming queries. A custom ExecutionPlan implementation was added specifically for the metadata cache that can report the predicates that are pushed down to the cache during query planning/execution. A big set of tests was added to to check that queries are working, and that predicates are being pushed down properly.	2024-11-22 10:57:26 -05:00
praveen-influx	33c2d47ba9	feat: drop/delete database (#25549 ) * feat: drop/delete database This commit allows soft deletion of database using `influxdb3 database delete <db_name>` command. The write buffer and last value cache are cleared as well. closes: https://github.com/influxdata/influxdb/issues/25523 * feat: reuse same code path when deleting database - In previous commit, the deletion of database immediately triggered clearing last cache and query buffer. But on restarts same logic had to be repeated to allow deleting database when starting up. This commit removes immediate deletion by explicitly calling necessary methods and moves the logic to `apply_catalog_batch` which already applies `CatalogOp` and also clearing cache and buffer in `buffer_ops` method which has hooks to call other places. closes: https://github.com/influxdata/influxdb/issues/25523 * feat: use reqwest query api for query param Co-authored-by: Trevor Hilton <thilton@influxdata.com> * feat: include deleted flag in DatabaseSnapshot - `DatabaseSchema` serialization/deserialization is delegated to `DatabaseSnapshot`, so the `deleted` flag should be included in `DatabaseSnapshot` as well. - insta test snapshots fixed closes: https://github.com/influxdata/influxdb/issues/25523 * feat: address PR comments + tidy ups --------- Co-authored-by: Trevor Hilton <thilton@influxdata.com>	2024-11-19 16:08:14 +00:00
Trevor Hilton	53f54a6845	feat: metadata cache core impl (#25552 ) * feat: core metadata cache structs with basic tests Implement the base MetaCache type that holds the hierarchical structure of the metadata cache providing methods to create and push rows from the WAL into the cache. Added a prune method as well as a method for gathering record batches from a meta cache. A test was added to check the latter for various predicates and that the former works, though, pruning shows that we need to modify how record batches are produced such that expired entries are not emitted. * refactor: filter expired entries and do some clean up in the meta cache	2024-11-18 12:28:12 -05:00
praveen-influx	814eb31309	chore: update core deps (#25532 ) * chore: update core deps - arrow/parquet deps are patched (as in core) - three specific code changes to cope with changes in core crates - TransitionPartitionId, use `from_parts` instead of `new` - arrow buffers can take &[u8] directly without `to_vec()`/`vec!` (used only in tests) - `schema` and `influxdb_line_protocol` crates need `v3` feature enabled * chore: update deny.toml * chore: formatting and deny toml changes Unicode-3.0 license is added to allowed licenses list, without it end up with 19 errors (`zerovec`, `zerovec-derive` etc.) * chore: address PR feedback - move enabling v3 feature to root Cargo.toml - added the upstream PR for datafusion-common that introduced RUSTSEC-2024-0384	2024-11-12 16:07:31 +00:00
Trevor Hilton	ec01934c57	chore: remove unnecessary rustsec for the tonic cve (#25516 ) `cargo deny` was showing that no crate matched the advisory criteria for this [RUSTSEC advisory](https://rustsec.org/advisories/RUSTSEC-2024-0376.html), so this PR removes the ignore entry. In addition, the `hashbrown` crate was causing a new audit failure, and updating it required that the `Zlib` license be added to our list of allowed licenses. No issue for this, but it is blocking another PR at the moment (https://github.com/influxdata/influxdb/pull/25515).	2024-11-04 15:02:39 -05:00
Trevor Hilton	d26a73802a	refactor: move to `ColumnId` and `Arc<str>` as much as possible (#25495 ) Closes #25461 _Note: the first three commits on this PR are from https://github.com/influxdata/influxdb/pull/25492_ This PR makes the switch from using names for columns to the use of `ColumnId`s. Where column names are used, they are represented as `Arc<str>`. This impacts most components of the system, and the result is a fairly sizeable change set. The area where the most refactoring was needed was in the last-n-value cache. One of the themes of this PR is to rely less on the arrow `Schema` for handling the column-level information, and tracking that info in our own `ColumnDefinition` type, which captures the `ColumnId`. I will summarize the various changes in the PR below, and also leave some comments in-line in the PR. ## Switch to `u32` for `ColumnId` The `ColumnId` now follows the `DbId` and `TableId`, and uses a globally unique `u32` to identify all columns in the database. This was a change from using a `u16` that was only unique within the column's table. This makes it easier to follow the patterns used for creating the other identifier types when dealing with columns, and should reduce the burden of having to manage the state of a table-scoped identifier. ## Changes in the WAL/Catalog * `WriteBatch` now contains no names for tables or columns and purely uses IDs * This PR relies on `IndexMap` for `_Id`-keyed maps so that the order of elements in the map is consistent. This has important implications, namely, that when iterating over an ID map, the elements therein will always be produced in the same order which allows us to make assertions on column order in a lot of our tests, and allows for the re-introduction of `insta` snapshots for serialization tests. This map type provides O(1) lookups, but also provides _fast_ iteration, which should help when serializing these maps in write batches to the WAL. * Removed the need to serialize the bi-directional maps for `DatabaseSchema`/`TableDefinition` via use of `SerdeVecMap` (see comments in-line) * The `tables` map in `DatabaseSchema` no stores an `Arc<TableDefinition>` so that the table definition can be shared around more easily. This meant that changes to tables in the catalog need to do a clone, but we were already having to do a clone for changes to the DB schema. * Removal of the `TableSchema` type and consolidation of its parts/functions directly onto `TableDefinition` * Added the `ColumnDefinition` type, which represents all we need to know about a column, and is used in place of the Arrow `Schema` for column-level meta-info. We were previously relying heavily on the `Schema` for iterating over columns, accessing data types, etc., but this gives us an API that we have more control over for our needs. The `Schema` is still held at the `TableDefinition` level, as it is needed for the query path, and is maintained to be consistent with what is contained in the `ColumnDefinition`s for a table. ## Changes in the Last-N-Value Cache * There is a bigger distinction between caches that have an explicit set of value columns, and those that accept new fields. The former should be more performant. * The Arrow `Schema` is managed differently now: it used to be updated more than it needed to be, and now is only updated when a row with new fields is pushed to a cache that accepts new fields. ## Changes in the write-path * When ingesting, during validation, field names are qualified to their associated column ID	2024-11-01 16:42:57 -04:00
Trevor Hilton	0e814f5d52	feat: SerdeVecMap type for serializing ID maps (#25492 ) This PR introduces a new type `SerdeVecHashMap` that can be used in places where we need a HashMap with the following properties: 1. When serialized, it is serialized as a list of key-value pairs, instead of a map 2. When deserialized, it assumes the serialization format from (1.) and deserializes from a list of key-value pairs to a map 3. Does not allow for duplicate keys on deserialization This is useful in places where we need to create map types that map from an identifier (integer) to some value, and need to serialize that data. For example: in the WAL when serializing write batches, and in the catalog when serializing the database/table schema. This PR refactors the code in `influxdb3_wal` and `influxdb3_catalog` to use the new type for maps that use `DbId` and `TableId` as the key. Follow on work can give the same treatment to `ColumnId` based maps once that is fully worked out. ## Explanation If we have a `HashMap<u32, String>`, `serde_json` will serialize it in the following way: ```json {"0": "foo", "1": "bar"} ``` i.e., the integer keys are serialized as strings, since JSON doesn't support any other type of key in maps. `SerdeVecHashMap<u32, String>` will be serialized by `serde_json` in the following way: ```json, [[0, "foo"], [1, "bar"]] ``` and will deserialize from that vector-based structure back to the map. This allows serialization/deserialization to run directly off of the `HashMap`'s `Iterator`/`FromIterator` implementations. ## The Controversial Part One thing I also did in this PR was switch the catalog from using a `BTreeMap` for tables to using the new `HashMap` type. This breaks the deterministic ordering of the database schema's `tables` map and therefore wrecks the snapshot tests we were using. I had to comment those parts of their respective tests out, because there isn't an easy way to make the underlying hashmap have a deterministic ordering just in tests that I am aware of. If we think that using `BTreeMap` in the catalog is okay over a `HashMap`, then I think it would be okay to roll a similar `SerdeVecBTreeMap` type specifically for the catalog. Coincidentally, this may actually be a good use case for [`indexmap`](https://docs.rs/indexmap/latest/indexmap/), since it holds supposedly similar lookup performance characteristics to hashmap, while preserving order and _having faster iteration_ which could be a win for WAL serialization speed. It also accepts different hashing algorithms so could be swapped in with FNV like `HashMap` can. ## Follow-up work Use the `SerdeVecHashMap` for column data in the WAL following https://github.com/influxdata/influxdb/issues/25461	2024-10-25 13:49:02 -04:00
praveen-influx	1f1125c767	refactor: udpate docs and tests for the telemetry crate (#25432 ) - Introduced traits, `ParquetMetrics` and `SystemInfoProvider` to enable writing easier tests - Uses mockito for code that depends on reqwest::Client and also uses mockall to generally mock any traits like `SystemInfoProvider` - Minor updates to docs	2024-10-08 15:45:13 +01:00
Trevor Hilton	42672e06b0	chore: cargo update (#25433 )	2024-10-07 10:50:53 -04:00
Michael Gattozzi	eeb1aa7905	feat: swap over to DbId and TableId everywhere (#25421 ) * feat: Add TableId and ColumnId * feat: swap over to DbId and TableId everywhere This commit swaps us over to using the DbId and TableId types everywhere for our internal systems. Anywhere that's external facing, such as names for last cache tables or line protocol parsing, use names. In these cases we have the `Catalog` which keeps a map of TableIds and DbIds in a bidirectional mapping for easy lookup i.e. id <-> names. While in essence the change itself isn't that complicated given the nature of how much we depended on names for things, the changes end up being quite invasive and extensive. Luckily it shouldn't be too hard to review. Note this does not add the column ids which will be done in a follow up PR. Closes #25375 Closes #25403 Closes #25404 Closes #25405 Closes #25412 Closes #25413	2024-10-03 14:47:46 -04:00
praveen-influx	8ccb580162	feat: telemetry report for parquet metrics (#25425 ) - added mechanism within PersistedFile to expose parquet file related metrics. The details are updated when new snapshot is generated and also when all snapshots are loaded when the process starts up - at the point of creating the telemetry payload these parquet metrics are looked up before sending it to the server. Closes: https://github.com/influxdata/influxdb/issues/25418	2024-10-03 15:11:40 +01:00
Trevor Hilton	7d37bbbce7	test: add test helpers for object store types (#25420 ) This adds a new crate `influxdb3_test_helpers` which provides two object store helper types that can be used to track request counts made through the store, as well as synchronize requests made through the store, resp.	2024-10-02 14:45:12 -04:00
praveen-influx	72dcd1866f	feat(telemetry): adds reads and writes (#25409 ) - instrumented code to get read and write measurement - introduced EventsBucket for collection of reads/writes - sampler now samples every minute for all metrics (including reads/writes) - other tidy ups closes: https://github.com/influxdata/influxdb/issues/25372	2024-10-01 18:34:00 +01:00
Trevor Hilton	4184a331ea	refactor: parquet cache with less locking (#25389 ) Closes #25382 Closes #25383 This refactors the parquet cache to use less locking by switching from using the `clru` crate to a hand-rolled cache implementation. The new cache still acts as an LRU, but it uses atomics to track hit-time per entry, and handles pruning in a separate process that is decoupled from insertion/gets to the cache. The `Cache` type uses a [`DashMap`](https://docs.rs/dashmap/latest/dashmap/struct.DashMap.html) internally to store cache entries. This should help reduce lock contention, and also has the added benefit of not requiring mutability to insert into _or_ get from the map. The cache maps an `object_store::Path` to a `CacheEntry`. On a hit, an entry will have its `hit_time` (an `AtomicI64`) incremented. During a prune operation, entries that have the oldest hit times will be removed from the cache. See the `Cache::prune` method for details. The cache is setup with a memory _capacity_ and a _prune percent_. The cache tracks memory used when entries are added, based on their _size_, and when a prune is invoked in the background, if the cache has exceeded its capacity, it will prune `prune_percent * cache.len()` entries from the cache. Two tests were added: * `cache_evicts_lru_when_full` to check LRU behaviour of the cache * `cache_hit_while_fetching` to check that a cache entry hit while a request is in flight to fetch that entry will not result in extra calls to the underlying object store	2024-09-27 11:59:17 -04:00
praveen-influx	29daa8332d	feat(telemetry): static values, cpu and mem metrics gathering (#25380 ) - basic setup to initialise the static values for telemetry store added. - cpu and memory used by influxdb3 is sampled at 1min interval - some minor tidyups Closes: https://github.com/influxdata/influxdb/issues/25370, https://github.com/influxdata/influxdb/issues/25371	2024-09-24 20:19:21 +01:00
Trevor Hilton	9c71b3ce25	feat: memory-cached object store for parquet files (#25377 ) Part of #25347 This sets up a new implementation of an in-memory parquet file cache in the `influxdb3_write` crate in the `parquet_cache.rs` module. This module introduces the following types: * `MemCachedObjectStore` - a wrapper around an `Arc<dyn ObjectStore>` that can serve GET-style requests to the store from an in-memory cache * `ParquetCacheOracle` - an interface (trait) that can accept requests to create new cache entries in the cache used by the `MemCachedObjectStore` * `MemCacheOracle` - implementation of the `ParquetCacheOracle` trait ## `MemCachedObjectStore` This takes inspiration from the [`MemCacheObjectStore` type](`1eaa4ed5ea/object_store_mem_cache/src/store.rs (L205-L213)`) in core, but has some different semantics around its implementation of the `ObjectStore` trait, and uses a different cache implementation. The reason for wrapping the object store is that this ensures that any GET-style request being made for a given object is served by the cache, e.g., metadata requests made by DataFusion. The internal cache comes from the [`clru` crate](https://crates.io/crates/clru), which provides a least-recently used (LRU) cache implementation that allows for weighted entries. The cache is initialized with a capacity and entries are given a weight on insert to the cache that represents how much of the allotted capacity they will take up. If there isn't enough room for a new entry on insert, then the LRU item will be removed. ### Limitations of `clru` The `clru` crate conveniently gives us an LRU eviction policy but its API may put some limitations on the system: * gets to the cache require an `&mut` reference, which means that the cache needs to be behind a `Mutex`. If this slows down requests through the object store, then we may need to explore alternatives. * we may want more sophisticated eviction policies than a straight LRU, i.e., to favour certain tables over others, or files that represent recent data over those that represent old data. ## `ParquetCacheOracle` / `MemCacheOracle` The cache oracle is responsible for handling cache requests, i.e., to fetch an item and store it in the cache. In this PR, the oracle runs a background task to handle these requests. I defined this as a trait/struct pair since the implementation may look different in Pro vs. OSS.	2024-09-24 10:58:15 -04:00
praveen-influx	c1a5e1b5fd	feat(telemetry): added basic types (#25374 ) - `TelemetryStore` is exposed for holding telemetry samples - added influxdb3_telemetry dependency to influxdb3 crate	2024-09-20 19:20:54 +01:00
Michael Gattozzi	54d209d0bf	feat: Add u32 ID for Databases (#25302 ) * feat: Remove lock for FileId tests Since we now are using cargo-nextest in CI we can remove the locks used in the FileId tests to make sure that we have no race conditions * feat: Add u32 ID for Databases This commit adds a new DbId for databases. It also updates paths to use that id as part of the name. When starting up the WriteBuffer we apply the DbId from the persisted snapshot much like we do for ParquetFileId's This introduces the influxdb3_id crate to avoid circular deps with ids. The ParquetFileId should also be moved into this crate, but it's outside the scope of this change. Closes #25301	2024-09-18 11:44:04 -04:00
praveen-influx	0c1fced7a4	refactor(catalog): catalog initialisation refactor (#25360 ) - when no catalog is found, create a new one with instance id and persist it immediately - enabled test-log in influxdb3_write Closes: https://github.com/influxdata/influxdb/issues/25346	2024-09-18 15:23:17 +01:00
praveen-influx	245b49ae0e	feat: add instance id and host id to catalog (#25343 ) - uses Arc<str> to represent create once and read everywhere type of string - updated snapshots for insta asserts, uses redaction to hardcode randomly generated UUID strings - added methods to catalog to expose instace and host ids Closes: https://github.com/influxdata/influxdb/issues/25315	2024-09-17 16:32:21 +01:00
praveen-influx	dd8f324728	chore: update core dependency revision (#25306 ) - version changed from 1d5011bde4c343890bb58aa77415b20cb900a4a8 to 1eaa4ed5ea147bc24db98d9686e457c124dfd5b7	2024-09-12 14:54:46 +01:00
Trevor Hilton	ad2ca83d72	chore: sync to latest core (#25284 )	2024-09-06 13:49:38 -04:00
dependabot[bot]	cb76f7a63c	chore(deps): bump quinn-proto from 0.11.6 to 0.11.8 (#25280 ) Bumps [quinn-proto](https://github.com/quinn-rs/quinn) from 0.11.6 to 0.11.8. - [Release notes](https://github.com/quinn-rs/quinn/releases) - [Commits](https://github.com/quinn-rs/quinn/compare/quinn-proto-0.11.6...quinn-proto-0.11.8) --- updated-dependencies: - dependency-name: quinn-proto dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2024-09-04 13:04:13 -04:00

1 2 3 4 5 ...

2751 Commits (praveen/reproduce-empty-snapshot-issue)