pgcat

mirror of https://github.com/postgresml/pgcat.git synced 2026-07-16 17:39:06 +00:00

Author	SHA1	Message	Date
Jose Fernández	58ce76d9b9	Refactor stats to use atomics (#375 ) * Refactor stats to use atomics When we are dealing with a high number of connections, generated stats cannot be consumed fast enough by the stats collector loop. This makes the stats subsystem inconsistent and a log of warning messages are thrown due to unregistered server/clients. This change refactors the stats subsystem so it uses atomics: - Now counters are handled using U64 atomics - Event system is dropped and averages are calculated using a loop every 15 seconds. - Now, instead of snapshots being generated ever second we keep track of servers/clients that have registered. Each pool/server/client has its own instance of the counter and makes changes directly, instead of adding an event that gets processed later. * Manually mplement Hash/Eq in `config::Address` ignoring stats * Add tests for client connection counters * Allow connecting to dockerized dev pgcat from the host * stats: Decrease cl_idle when idle socket disconnects	2023-03-28 17:19:37 +02:00
Zain Kabani	ca4431b67e	Add idle client in transaction configuration (#380 ) * Add idle client in transaction configuration * fmt * Update docs * trigger build * Add tests * Make the config dynamic from reloads * fmt * comments * trigger build * fix config.md * remove error	2023-03-24 08:20:30 -07:00
Lev Kokotov	b4baa86e8a	Extended query protocol sharding (#339 ) * Prepared stmt sharding s tests * len check * remove python test * latest rust * move that to debug for sure * Add the actual tests * latest image * Update tests/ruby/sharding_spec.rb	2023-03-10 07:55:22 -08:00
Mostafa Abdelraouf	aa89e357e0	PgCat Query Mirroring (#341 ) This is an implementation of Query mirroring in PgCat (outlined here #302) In configs, we match mirror hosts with the servers handling the traffic. A mirror host will receive the same protocol messages as the main server it was matched with. This is done by creating an async task for each mirror server, it communicates with the main server through two channels, one for the protocol messages and one for the exit signal. The mirror server sends the protocol packets to the underlying PostgreSQL server. We receive from the underlying PostgreSQL server as soon as the data is available and we immediately discard it. We use bb8 to manage the life cycle of the connection, not for pooling since each mirror server handler is more or less single-threaded. We don't have any connection pooling in the mirrors. Matching each mirror connection to an actual server connection guarantees that we will not have more connections to any of the mirrors than the parent pool would allow.	2023-03-10 06:23:51 -06:00
Mostafa Abdelraouf	2cc6a09fba	Add Manual host banning to PgCat (#340 ) Sometimes we want an admin to be able to ban a host for some time to route traffic away from that host for reasons like partial outages, replication lag, and scheduled maintenance. We can achieve this today using a configuration update but a quicker approach is to send a control command to PgCat that bans the replica for some specified duration. This command does not change the current banning rules like Primaries cannot be banned When all replicas are banned, all replicas are unbanned	2023-03-06 06:10:59 -06:00
Jose Fernández	8a0da10a87	Dev environment (#338 ) Add dev env	2023-03-02 12:14:10 -05:00
zainkabani	eb8cfdb1f1	Adds SHUTDOWN command as alternate option to sending SIGINT (#331 ) * Adds SHUTDOWN command to PgCat as alternate option to sending SIGINT * Check if we're already in SHUTDOWN sequence * Send signal directly from shutdown instead of using channel * Add tests * trigger build * Lowercase response and boolean change * Update tests * Fix tests * typo	2023-02-26 22:16:30 -08:00
Mostafa Abdelraouf	75a7d4409a	Fix Back-and-forth RELOAD Bug (#330 ) We identified a bug where RELOAD fails to update the pools. To reproduce you need to start at some config state, modify that state a bit, reload, revert the configs back to the original state, and reload. The last reload will fail to update the pool because PgCat "thinks" the pool state didn't change. This is because we use a HashSet to keep track of config hashes but we never remove values from it. Say we start with State A, we modify pool configs to State B and reload. Now the POOL_HASHES struct has State A and State B. Attempting to go back to State A will encounter a hashset hit which is interpreted by PgCat as "Configs are the same, no need to reload pools" We fix this by attaching a config_hash value to ConnectionPool object and we calculate that value when we create the pool. This eliminates the need for a global variable. One shortcoming here is that changing any config under one user in the pool will trigger a reload for the entire pool (which is fine I think)	2023-02-21 21:53:10 -06:00
Nicholas Dujay	37e1c5297a	implement show users (#329 ) * implement show users * fix compile errors * add basic ruby test * gitignore things	2023-02-21 13:08:43 -08:00
Mostafa Abdelraouf	28f2d19cac	More coverage cleanup (#328 ) Apply a new style + remove function coverage report	2023-02-17 09:18:54 -06:00
Mostafa Abdelraouf	f9134807d7	More Test coverage + fix some code coverage bugs (#321 ) Connection to the CI databases is viewed by Postgres as coming from localhost. The pg_hba.conf file generated by the docker image uses trust for these connections, that's why we had no test coverage on SASL and md5 branches. This PR fixes this issue. There was also an issue with under-reporting code coverage. This should be fixed now	2023-02-16 23:09:22 -06:00
Mostafa Abdelraouf	bf6efde8cc	Fix code coverage + less flakiness (#318 ) Code coverage logic was missing coverage from rust tests. This is now fixed. Also, we weren't reaping spawned PgCat processes correctly which left zombie processes.	2023-02-13 15:29:08 -06:00
Mostafa Abdelraouf	f1265a5570	Introduce tcp_keepalives to PgCat (#315 ) We have encountered a case where PgCat pools were stuck following a database incident. Our best understanding at this point is that the PgCat -> Postgres connections died silently and because Tokio defaults to disabling keepalives, connections in the pool were marked as busy forever. Only when we deployed PgCat did we see recovery. This PR introduces tcp_keepalives to PgCat. This sets the defaults to be keepalives_idle: 5 # seconds keepalives_interval: 5 # seconds keepalives_count: 5 # a count These settings can detect the death of an idle connection within 30 seconds of its death. Please note that the connection can remain idle forever (from an application perspective) as long as the keepalive packets are flowing so disconnection will only occur if the other end is not acknowledging keepalive packets (keepalive packet acks are handled by the OS, the application does not need to do anything). I plan to add tcp_user_timeout in a follow-up PR.	2023-02-08 11:35:38 -06:00
Mostafa Abdelraouf	87a771aecc	Log error messages for network failures (#289 ) We are seeing some Error reading message code from socket error messages, we want to get more context so this PR logs the actual error reported.	2023-01-19 05:18:08 -06:00
dependabot[bot]	99a3b9896d	chore(deps): bump activerecord from 7.0.3.1 to 7.0.4.1 in /tests/ruby (#287 ) Bumps [activerecord](https://github.com/rails/rails) from 7.0.3.1 to 7.0.4.1. - [Release notes](https://github.com/rails/rails/releases) - [Changelog](https://github.com/rails/rails/blob/v7.0.4.1/activerecord/CHANGELOG.md) - [Commits](https://github.com/rails/rails/compare/v7.0.3.1...v7.0.4.1) --- updated-dependencies: - dependency-name: activerecord dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2023-01-18 16:56:16 -08:00
Mostafa Abdelraouf	7894bba59b	Introduce least-outstanding-connections load balancing (#282 ) Least outstanding connections load balancing can improve the load distribution between instances but for Pgcat it may also improve handling slow replicas that don't go completely down. With LoC, traffic will quickly move away from the slow replica without waiting for the replica to be banned. If all replicas slow down equally (due to a bad query that is hitting all replicas), the algorithm will degenerate to Random Load Balancing (which is what we had in Pgcat until today). This may also allow Pgcat to accommodate pools with differently-sized replicas.	2023-01-17 06:52:18 -06:00
zainkabani	19f635881a	Don't send discard all when state is changed in transaction (#186 ) * Don't send discard all when state is changed in transaction * Remove unnecessary clone * spelling * Move transaction check to SET command * Add test for set command in transaction * type * Update comments * Update comments * use moves instead of clones for initial message * don't make message mutable * Update unwrap * but i'm not a wrapper * Add set local test * change continue	2022-10-13 19:33:12 -07:00
Mostafa Abdelraouf	3d33ccf4b0	Fix maxwait metric (#183 ) Max wait was being reported as 0 after #159 This PR fixes that and adds test	2022-10-05 21:41:09 -05:00
Mostafa Abdelraouf	af064ef447	Set client state to idle after error (#179 ) * Set client state to idle after error * fmt * spelling * clean up	2022-09-24 09:09:15 -07:00
Mostafa Abdelraouf	f7a951745c	Report Query times (#166 ) * Report avg and total query timing * Report query times * fmt	2022-09-15 02:21:45 -04:00
Mostafa Abdelraouf	4ae1bc8d32	Add SHOW CLIENTS / SHOW SERVERS + Stats refactor and tests (#159 ) * wip * Main Thread Panic when swarmed with clients * fix * fix * 1024 * fix * remove test * Add SHOW CLIENTS * revert * fmt * Refactor + tests * fmt * add test * Add SHOW SERVERS + Make PR unreviewable * prometheus * add state to clients and servers * fmt * Add application_name to server stats * Add tests for waiting clients * Docs * remove comment * comments * typo * cleanup * CI	2022-09-14 11:20:41 -04:00
Mostafa Abdelraouf	9514b3b2d1	Clean connection state up after protocol named prepared statement (#163 ) * Clean connection state up after protocol named prepared statement * Avoid cloning + add test * fmt	2022-09-07 20:37:17 -07:00
Lev Kokotov	6d41640ea9	Send signal even if process is gone (#162 ) * Send signal even if process is gone * hmm * hmm	2022-09-07 09:22:52 -07:00
Mostafa Abdelraouf	23a642f4a4	Send DISCARD ALL even if client is not in transaction (#152 ) * Send DISCARD ALL even if client is not in transaction * fmt * Added tests + avoided sending extra discard all * Adds set name logic to beginning of handle client * fmt * refactor dead code handling * Refactor reading command tag * remove unnecessary trim * Removing debugging statement * typo * typo{ * documentation * edit text * un-unwrap * run ci * run ci Co-authored-by: Zain Kabani <zain.kabani@instacart.com>	2022-09-01 20:06:55 -07:00
Mostafa Abdelraouf	65b69b46d2	Allow running integration tests with coverage locally (#151 )	2022-08-30 10:43:45 -07:00
Mostafa Abdelraouf	d48c04a7fb	Ruby integration tests (#147 ) * Ruby integration tests * forgot a file * refactor * refactoring * more refactoring * remove config helper * try multiple databases * fix * more databases * Use pg stats * ports * speed * Fix tests * preload library * comment	2022-08-30 09:14:53 -07:00
Lev Kokotov	9d84d6f131	Graceful shutdown and refactor (#144 ) * Graceful shutdown and refactor * ok * _Graceful_ shutdown * Remove hardcoded setting * clean up * end * timeout * hmm * hmm! * bash * bash * hmm * maybe maybe * Adds tests and move non-admin connection rejection to startup (#145) * Move error response * Adds tests and removes unused variable * Adds debug log Co-authored-by: zainkabani <77307340+zainkabani@users.noreply.github.com>	2022-08-25 06:40:56 -07:00
Mostafa Abdelraouf	c054ff068d	Avoid sending `Z` packet in the middle of extended protocol packet sequence if we fail to get connection from pool (#137 ) * Failing test * maybe * try fail * try * add message * pool size * correct user * more * debug * try fix * see stdout * stick? * fix configs * modify * types * m * maybe * make tests idempotent * hopefully fails * Add client fix * revert pgcat.toml change * Fix tests	2022-08-23 11:02:23 -07:00
zainkabani	65c32ad9fb	Validates pgcat is closed after shutdown python tests (#116 ) * Validates pgcat is closed after shutdown python tests * Fix pgrep logic * Moves sigterm step to after cleanup to decouple * Replace subprocess with os.system for running pgcat	2022-08-09 14:09:53 -07:00
Mostafa Abdelraouf	7592339092	Prevent clients from sticking to old pools after config update (#113 ) * Re-acquire pool at the beginning of Protocol loop * Fix query router + add tests for recycling behavior	2022-08-09 12:18:27 -07:00
zainkabani	3719c22322	Implementing graceful shutdown (#105 ) * Initial commit for graceful shutdown * fmt * Add .vscode to gitignore * Updates shutdown logic to use channels * fmt * fmt * Adds shutdown timeout * Fmt and updates tomls * Updates readme * fmt and updates log levels * Update python tests to test shutdown * merge changes * Rename listener rx and update bash to be in line with master * Update python test bash script ordering * Adds error response message before shutdown * Add details on shutdown event loop * Fixes response length for error * Adds handler for sigterm * Uses ready for query function and fixes number of bytes * fmt	2022-08-08 16:01:24 -07:00
Mostafa Abdelraouf	106ebee71c	Fix local dev (#112 ) * Fix Dev env * Update tests/sharding/query_routing_setup.sql * Update tests/sharding/query_routing_setup.sql * bring pgcat.toml on ci and local dev to parity * more parity * pool names * pool names * less diff * fix tests * fmt * add other user to setup Co-authored-by: Lev Kokotov <levkk@users.noreply.github.com>	2022-08-08 13:15:48 -07:00
Mostafa Abdelraouf	5ac85eaadd	Fix Python tests and remove CircleCI-specific path (#106 ) * Remove CircleCI-specific path in tests * ..? * Fix testsP * Fix python test * remove pip * Maybe fail? * return code? * no & * Fix tests	2022-08-02 15:52:22 -07:00
Mostafa Abdelraouf	1b648ca00e	Send proper server parameters to clients using admin db (#103 ) * Send proper server parameters to clients using admin db * clean up * fix python test * build * Add python * missing & * debug ls * fix tests * fix tests * fix * Fix warning * Address comments	2022-07-31 19:52:23 -07:00
Mostafa Abdelraouf	2ae4b438e3	Add support for multi-database / multi-user pools (#96 ) * Add support for multi-database / multi-user pools * Nothing * cargo fmt * CI * remove test users * rename pool * Update tests to use admin user/pass * more fixes * Revert bad change * Use PGDATABASE env var * send server info in case of admin	2022-07-27 19:47:55 -07:00
dependabot[bot]	eff8e3e229	Bump activerecord from 7.0.2.2 to 7.0.3.1 in /tests/ruby (#94 ) Bumps [activerecord](https://github.com/rails/rails) from 7.0.2.2 to 7.0.3.1. - [Release notes](https://github.com/rails/rails/releases) - [Changelog](https://github.com/rails/rails/blob/v7.0.3.1/activerecord/CHANGELOG.md) - [Commits](https://github.com/rails/rails/compare/v7.0.2.2...v7.0.3.1) --- updated-dependencies: - dependency-name: activerecord dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2022-07-12 13:24:41 -07:00
Lev Kokotov	b93303eb83	Live reloading entire config and bug fixes (#84 ) * Support reloading the entire config (including sharding logic) without restart. * Fix bug incorrectly handing error reporting when the shard is set incorrectly via SET SHARD TO command. selected wrong shard and the connection keep reporting fatal #80. * Fix total_received and avg_recv admin database statistics. * Enabling the query parser by default. * More tests.	2022-06-24 14:52:38 -07:00
Lev Kokotov	37e3a86881	Pass application_name to server (#73 ) * Pass application_name to server * fmt	2022-06-03 00:15:50 -07:00
Lev Kokotov	ccbca66e7a	Poorly behaved client fix (#65 ) * Poorly behaved client fix * yes officer * fix tests * no useless rescue * Looks ok	2022-05-09 09:09:22 -07:00
Lev Kokotov	35828a0a8c	Per-shard statistics (#57 ) * per shard stats * aight * cleaner * fix show lists * comments * more friendly * case-insensitive * test all shards * ok * HUH?	2022-03-04 17:04:27 -08:00
Lev Kokotov	303fec063b	Ruby (#30 ) * cop * log	2022-02-20 23:33:04 -08:00
Lev Kokotov	a556ec1c43	More query router commands; settings last until changed again; docs (#25 ) * readme * touch up docs * stuff * refactor query router * remove unused * less verbose * docs * no link * method rename	2022-02-19 08:57:24 -08:00
Lev Kokotov	bbacb9cf01	Explicit shard selection; Rails tests (#24 ) * Explicit shard selection; Rails tests * try running ruby tests * try without lockfile * aha * ok	2022-02-18 09:43:07 -08:00
Lev Kokotov	4c8a3987fe	Refactor query routing into its own module (#22 ) * Refactor query routing into its own module * commments; tests; dead code * error message * safer startup * hm * dont have to be public * wow * fix ci * ok * nl * no more silent errors	2022-02-16 22:52:11 -08:00
Lev Kokotov	c1476d29da	config tests	2022-02-10 09:07:10 -08:00
Lev Kokotov	daf120aeac	more tests	2022-02-10 08:35:25 -08:00
Lev Kokotov	fccfb40258	nl	2022-02-09 21:20:20 -08:00
Lev Kokotov	a9b2a41a9b	fixes to the banlist	2022-02-09 21:19:14 -08:00
Lev Kokotov	dfc05c3dca	sharding readme	2022-02-08 17:15:35 -08:00
Lev Kokotov	9657256adf	fixed health check; sharding setup and tests	2022-02-08 15:48:28 -08:00

1 2

55 Commits