Compare commits

...
96 Commits
Author SHA1 Message Date
thebestvibedev 18c02a6bd6 Update stock/hsx.vn.md 2026-07-08 14:20:19 +07:00
thebestvibedev 253b7b3de5 Đầu tư chứng cho mau giàu 2026-07-08 14:20:00 +07:00
thebestvibedev dbe87937a0 Trộm tiếp 2026-07-08 13:53:49 +07:00
thebestvibedev 322f06d77e Trộm chó hahaha 2026-07-08 13:53:01 +07:00
Vitaliy Filippov 5dbc679e16 Fix padded block checksums - v2 2026-07-05 14:09:57 +03:00
Vitaliy Filippov d48864a6be Fix unaligned pointer warnings 2026-07-04 21:40:29 +03:00
Vitaliy Filippov 462482d319 Fix zero-padded big_write checksum verification in the new store 2026-07-04 21:37:22 +03:00
Vitaliy Filippov 5ef9d78461 Release 3.0.15
- QEMU virtual disk migration with enabled iothread is finally fixed correctly.
- Fixed operation of Proxmox VMs with swTPM without enabling NBD for all disks.
- Debian packages are now again built with stable Antietcd released instead of
  the unstable master branch.
- Antietcd cluster mode previously broken in that master branch is fixed. The
  symptom was Antietcd being unable to elect the leader in a cluster.
- Fixed monitor startup with embedded Antietcd when using IPv6.
- Fixed space statistics calculation for FS and S3 pools in the new storage ([PR #127](https://github.com/vitalif/vitastor/pull/127)).
2026-06-28 11:29:38 +03:00
Vitaliy Filippov 1b20010a69 Install npm packages in debian/rules 2026-06-28 11:29:27 +03:00
Vitaliy Filippov d750d00c6f Fix SWTPM in Proxmox 2026-06-27 16:04:51 +03:00
Vitaliy Filippov 66cb564e2c Fix background jobs in etcd_fail test 2026-06-27 15:42:53 +03:00
Vitaliy Filippov f782decbd6 Update antietcd to 1.3.1 2026-06-27 15:42:42 +03:00
Vitaliy Filippov 717582c4ba Use release builds of antietcd&tinyraft in debian packages (not master branch) 2026-06-27 10:27:15 +03:00
Vitaliy Filippov cd62c0755f Fix IPv6 antietcd configuration 2026-06-27 00:16:03 +03:00
c6f5ab7d79 fix pool stats with no_inode_stats (#127)
Co-authored-by: zhu.chengzhen <zhu.chengzhen@jingjiamicro.com>
2026-06-24 12:41:56 +03:00
Vitaliy Filippov 09607ddcbe Actual fix for live migration with iothread 2026-06-23 00:07:34 +03:00
Vitaliy Filippov 278852b4d5 Release 3.0.14
How many (bugs!) I've knifed, how many I've slit!

General note: most bug fixes now include regression tests to verify that they don't repeat in the future.
Most bugs fixed in this release were detected by using LLM analysis (Claude Opus/Fable, GPT 5.5).

OSD:
- Fix OSD hanging with an infinite loop when setting autosync_interval to 0 at runtime
- (IMPORTANT) Fix EC PGs hanging in REPEERING when the last final commit/rollback in a
  batch completes with an error
- Limit pg_size by 64 because peering doesn't handle larger values with EC — they just
  lead to 'incomplete' objects
- Fix OSD crash with an "assertion failed" error on EIO retry in snapshot chain read
  (i.e. when some chunks belong to a corrupted replica with checksum mismatch)
- (IMPORTANT) Disable chunked PG count resharding due to possible interference with
  compaction changes (will be re-enabled after fixes)
- (IMPORTANT) Fix incorrect snapshot allocation bitmap recovery during EC chained read
- Add on-wire request size validation to prevent possible OOM/DoS/heap corruption
  on receiving invalid data from the network
- (IMPORTANT) Fix parity-less EC writes destroying snapshot allocation bitmaps
  (i.e. when all parity OSDs in a PG are missing)
- (IMPORTANT) Fix EC N+K, K>=2 recovery destroying snapshot allocation bitmaps
  of live parity chunks
- (IMPORTANT) Fix a possible OSD crash during EC misplaced object scrubbing
- (IMPORTANT) Verify object bitmap consistency during scrub (only data was checked previously)
- (IMPORTANT) Fix corrupted object chunks incorrectly marked as non-corrupted on the second scrub
- (IMPORTANT) Fix cached EC decoding of multiple stripes with ISA-L (ISA-L is the default)

New store:
- Fix a theoretically possible OSD crash on startup when using the previously added
  workaround for the "double-claim" problem
- Remove theoretically possible incorrect metadata block writes during batch EC COMMITs
  restarted due to a full metadata area
- Fix incorrect compaction counter tracking after OSD restart (could probably lead to
  compaction not restarted correctly after a restart)
- (IMPORTANT) Fix some of parallel big_writes possibly not waiting for data fsync, thus not providing durability
- Fix possible OSD crash on sync retry when io_uring is full
- Fix a possible crash during startup on corrupted on-disk data with too small entry sizes

Old store:
- Prevent loading extra garbage metadata entries from the last 4 MB of metadata area
- Fix read operations possibly crashing if a metadata read (with inmemory_metadata=false)
  was restarted due to a full io_uring
- Fix a possible memory leak of temporary buffers and bitmaps/checksums when a read
  was restarted due to a full io_uring (reproducible with either inmemory_metadata=false or block_size>256k)
- Fix a possible OSD crash during padded checksum reads if buffer count exceeded 1024
  (IOV_MAX) (reproducible only with csum_block_size > 4k and block_size >= 4M)
- (IMPORTANT) Fix partial padded read journal checksum verification with csum_block_size > 4k
- Fix incorrect marking of corrupted objects as non-corrupted after flushing data
  from journal (with inmemory_journal=false)
- (IMPORTANT) Fix deferred freeing of a different block when a block was used by a parallel read
- Fix per-inode statistics not being disabled for FS and S3 pools correctly, leading to etcd
  overload with unneeded per-inode statistics, slower etcd operation, increased memory usage,
  and too many Prometheus statistics exported by the monitor

Both stores:
- Fix possibly left garbage in the metadata area if the first OSD startup was interrupted —
  metadata header is now written only after initializating metadata
- Check for short reads during initialization (just in case, doesn't happen in real life)

Clients:
- Fix write-back queue item split in case when write-back is enabled at runtime
- Implement bdrv_detach_aio_context & bdrv_attach_aio_context in the QEMU driver (should fix migration with iothread)
- Do not crash on full io_uring in ublk server
- Fix missing --readonly option handling in NBD server
- Stop gracefully on NBD_CMD_DISC instead of just exit(0) in NBD server
- Fix writeback detection in ublk server for --image mode
- Limit the amount of incoming data for NFS clients to prevent choking on memory in async mount mode

Tools (vitastor-disk/vitastor-cli):
- Prevent vitastor-cli merge possibly exiting before completing the last sync/delete operations
- Fix vitastor-disk incorrectly validating too large small_write entry length
- Fix vitastor-cli merge ignoring input option validation errors
- Fix vitastor-cli rm-data always skipping the final fsync
- Fix vitastor-disk resize not moving the last used data block
- Fix vitastor-disk write-meta incorrectly importing new store small_write entries
- Fix vitastor-disk write-journal and write-meta importing old store data incorrectly
  when csum_block_size is > 4k
- Support --io option for vitastor-disk dump-journal/write-journal
- Fix vitastor-disk resize crash when converting from very old (0.5.x) metadata
- Fix vitastor-disk trim incorrectly rounding block ranges with --discard_granularity
  option explicitly set to a value > 4k, possibly leading to discarding live data
- Fix vitastor-disk write-meta importing new store metadata incorrectly with > 4 GB metadata area size
- (IMPORTANT) Fix vitastor-cli modify --resize to a smaller size clearing all image data O_o

Other:
- Do not crash with an uncaught exception when an invalid /osd/state/ with a non-numeric
  suffix is present in etcd (in OSD and all client services)
- Fix possible crash in vitastor-kv when handling a corrupted DB due to a uint32 overflow
- Fix NFS-RDMA memory allocator crashing in some situations
- Fix small shared file extend-write potentially reading unallocated memory (NFS)
- Add bounds checks to prevent uint32 overflows in NFS/XDR
- Re-enable accidentally disabled safety checks (asserts) in files with included cpp-btree
- Fix too small memory allocation in NFS portmap
2026-06-21 18:17:02 +03:00
Vitaliy Filippov 7b11c6e90d Add a regression test for v1 store attempting to load overflowing garbage from the end of metadata area (commit 6b003bcc34) 2026-06-21 17:23:50 +03:00
Vitaliy Filippov 6251ce8b9a Add a regression test for incorrect deferred freeing of data block (commit 43aa4cfff6) 2026-06-21 16:49:38 +03:00
Vitaliy Filippov d4a42f61cf Handle full io_uring in ublk-server 2026-06-21 01:40:49 +03:00
Vitaliy Filippov 6b003bcc34 Protect against done_cnt > block_count, just in case 2026-06-20 21:05:06 +03:00
Vitaliy Filippov 1c945bcb41 Make check_completed a macro in test_cluster_client 2026-06-20 20:18:31 +03:00
Vitaliy Filippov 716527b184 Fix missing --readonly handling in NBD server 2026-06-20 11:52:29 +03:00
Vitaliy Filippov 8c0486bd76 Fix bounds check in kv_db (fix invalid input handling?) 2026-06-20 11:43:33 +03:00
Vitaliy Filippov d2cf271f64 Track in_flight for sync & delete in vitastor-cli merge 2026-06-20 11:43:33 +03:00
Vitaliy Filippov d93b488e32 Just in case - OP_SYNC cannot fail but handle its error in vitastor-cli dd 2026-06-20 11:43:33 +03:00
Vitaliy Filippov fbffec5abb Add a regression test for the last 2 fixed bugs 2026-06-20 11:18:09 +03:00
Vitaliy Filippov b6eb8f2055 Add missing sqe retries on metadata read 2026-06-20 11:15:38 +03:00
Vitaliy Filippov e373ea2163 Fix freeing of dyn_data and temp metadata block buffers on read cancel
A regression test would also be fun
2026-06-20 02:10:11 +03:00
Vitaliy Filippov 5e12b4a1a5 Fix OOB read in old store checksum read if block count exceeds IOV_MAX (almost unreachable)
May be fun to write a regression test for it :) IOV_MAX is 1024, so it requires at least a 4 MB object...
2026-06-20 02:03:45 +03:00
Vitaliy Filippov 78b067566f Do not use std::stoull as it may throw 2026-06-20 01:54:49 +03:00
Vitaliy Filippov cf3abdb9e3 Clear autosync_timer when autosync_interval is set to 0
Clear other timers similarly (but their interval can't be 0)
2026-06-20 01:50:44 +03:00
Vitaliy Filippov 826b35b369 Do not try to erase_double_claim entries during iteration 2026-06-19 22:05:44 +03:00
Vitaliy Filippov cdc730314b Fix EC PGs hanging in REPEERING when flush error is last in the batch 2026-06-19 01:49:26 +03:00
Vitaliy Filippov 5248d7f324 Fix 2 bugs in NFS-RDMA allocator, add a test for both
1) alloc() could add the region start into freelist instead of its free part
2) free() was merging freed buffers incorrectly both forward and backward
2026-06-19 02:08:41 +03:00
Vitaliy Filippov 4161f0bd01 Remove doubtful metadata block write logic from blockstore_stable on metadata ENOSPC 2026-06-18 00:55:57 +03:00
Vitaliy Filippov 5627977a9b Limit pg_size by 64 2026-06-18 00:24:48 +03:00
Vitaliy Filippov ec9cfa76c1 Fix to_compact_count tracking on load in the new store 2026-06-18 00:21:03 +03:00
Vitaliy Filippov 0c88884576 Fix incorrect batch big_write fsync logic in the new store 2026-06-18 00:04:15 +03:00
Vitaliy Filippov ba7637d9ad Add a test for incorrect batch big_write fsync logic in the new store 2026-06-18 00:04:15 +03:00
Vitaliy Filippov ac1025c7a5 Fix read small_write.len before init 2026-06-17 17:41:51 +03:00
Vitaliy Filippov 2a81cef78a Fix crash on EIO retry in chained read 2026-06-17 16:38:20 +03:00
Vitaliy Filippov bcf6a7c7d1 Fix write-back queued items split in the client 2026-06-17 16:15:39 +03:00
Vitaliy Filippov 85c4be3957 Clear lsn on sync retry to prevent possible crash on full io_uring 2026-06-17 15:40:05 +03:00
Vitaliy Filippov cf2ba05e4b Disable chunked resharding -- it may interfere with journal flushing; will be re-enabled after fixes 2026-06-17 15:35:19 +03:00
Vitaliy Filippov 04c2f8d408 Stop gracefully on NBD_CMD_DISC instead of just exit(0) 2026-06-17 12:46:03 +03:00
Vitaliy Filippov 0e528ca8f3 Fix missing start_merge() error check 2026-06-17 12:42:23 +03:00
Vitaliy Filippov c6f733b96a Fix final sync in vitastor-cli rm-data 2026-06-17 12:41:22 +03:00
Vitaliy Filippov 938ac09248 Fix small shared file extend-write potentially reading unallocated memory 2026-06-17 12:36:04 +03:00
Vitaliy Filippov 3becdbf5b9 Add bounds checks to prevent uint32 overflows in NFS/XDR 2026-06-17 12:04:50 +03:00
Vitaliy Filippov 9c49315fdf Fix moving of the last block in vitastor-disk resize 2026-06-17 01:19:12 +03:00
Vitaliy Filippov 430d3cfb6f Add a dump|load test, fix multiple bugs in both new&old dump/load utils
Details:
- New store dump/write-meta didn't use actual metadata parameters from the header
- New store write-meta calculated small entry sizes incorrectly
- Old store write-journal imported entries with checksums incorrectly
- Old store write-meta didn't import block_csums at all (it was using a wrong json key)
2026-06-17 01:19:12 +03:00
Vitaliy Filippov c491db699c Implement bdrv_detach_aio_context & bdrv_attach_aio_context (should fix migration with iothread) 2026-06-17 00:39:28 +03:00
Vitaliy Filippov e9d053e30f Support --io option for vitastor-disk dump-journal/write-journal 2026-06-17 00:39:28 +03:00
Vitaliy Filippov 1be51f903c Check entry sizes in blockstore_heap during loading 2026-06-16 20:34:52 +03:00
Vitaliy Filippov cd51f14a90 Fix incorrect chained_read bitmap recovery 2026-06-15 21:10:44 +03:00
Vitaliy Filippov 9b8107875f Add a regression test for incorrect chained_read bitmap recovery 2026-06-15 21:10:44 +03:00
Vitaliy Filippov ca606570f7 Add request size validation 2026-06-15 01:56:54 +03:00
Vitaliy Filippov 22d094ccc6 Add a regression test to check that parity-less EC writes do not destroy bitmaps 2026-06-15 01:56:54 +03:00
Vitaliy Filippov fd84d84279 Add a regression test for scrub bitmap comparison (EC 3+3, only chunk 1 written and lost) 2026-06-15 01:52:04 +03:00
Vitaliy Filippov af2b1e28e3 Add a regression test for EC bitmap recovery losing a parity chunk 2026-06-15 01:52:04 +03:00
Vitaliy Filippov 9dda449f48 Fix fully-degraded EC writes destroying bitmaps 2026-06-15 01:52:04 +03:00
Vitaliy Filippov e6d4b32629 Also check padded small_write 2026-06-15 01:31:19 +03:00
Vitaliy Filippov 4de22a08e2 Fix partial padded read checksum verification 2026-06-15 01:31:19 +03:00
Vitaliy Filippov a403de46b3 Test partial padded read 2026-06-15 01:31:19 +03:00
Vitaliy Filippov ac20f605f6 Fix scrub bitmap comparison 2026-06-15 01:31:19 +03:00
Vitaliy Filippov 51ecbadb12 Only update bitmaps when writing data in primary subops
Fixes fully degraded EC writes and recovering EC writes corrupting bitmaps
2026-06-15 01:31:19 +03:00
Vitaliy Filippov 6a4627b625 Write a test for store v1 & journal corruption preservation, fix MULTIPLE bugs 2026-06-14 11:23:36 +03:00
Vitaliy Filippov fd2b8b8792 Fix ec_check_combination more correctly (fixes osd_rmw_test_je) 2026-06-13 19:56:06 +03:00
Vitaliy Filippov f3d662bac7 Verify bitmaps during scrub 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 553cc8ef87 Add a test for replicated bitmap scrub 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 67fdf1142b Fix vitastor-disk crash when converting from 0.5.x metadata 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 02f6e564a6 Old store - write header on init only after clearing metadata 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 33d14061d6 New store - write header on init only after clearing metadata 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 677e755e4e Replace VLA with a vector and also fix invalid dangling pointer warnings 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 11a972cbfb Delete pg.peering_state when it is not needed 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 1834743a0e Fix incorrect clearing of LOC_CORRUPTED on the second scrub 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 213f76c66c Add a regression test for LOC_CORRUPTED cleared on second scrub 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 91698404a7 Add a basic OSD test as an example 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 27bd38d95e Allow to create OSD with mocked blockstore and network 2026-06-13 19:56:06 +03:00
Vitaliy Filippov ef0e61be1b Split and mock etcd_state_client_t for testing 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 26fb08d7da Change st_cli to pointer 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 80fa3094b3 Fix cpp-btree disabling asserts everywhere it is included to! 2026-06-13 19:56:06 +03:00
Vitaliy Filippov df931b1e17 Check for short reads in blockstore init 2026-06-13 19:56:06 +03:00
Vitaliy Filippov 906294ae9a Fix ec_check_combination() short tmp_buf allocation 2026-06-13 19:56:05 +03:00
Vitaliy Filippov 15f69719e4 Not an actual bug, clean up fsync conditions for clarity
It could be a bug if data_fd could be equal to journal_fd but not to meta_fd,
but it can't.
2026-06-13 19:55:47 +03:00
Vitaliy Filippov 43aa4cfff6 Fix incorrect deferred freeing of block used by a parallel read in the old store 2026-06-13 19:55:47 +03:00
Vitaliy Filippov 6901227390 Fix cached EC decoding of multiple stripes with ISA-L 2026-06-10 02:07:58 +03:00
Vitaliy Filippov 7f718feaf6 Fix alloc in nfs_portmap 2026-06-10 01:02:30 +03:00
Vitaliy Filippov 236ffbb24e Fix multilist_alloc_t bug (not triggerable in real operation but still a bug) 2026-06-10 01:01:20 +03:00
Vitaliy Filippov 0e300f4c50 Fix incorrect rounding in vitastor-disk trim when --discard_granularity option is passed 2026-06-10 00:38:34 +03:00
Vitaliy Filippov 6d82a3daa3 Fix fill_block_empty_space argument type (ui32 -> ui64) 2026-06-10 00:33:31 +03:00
Vitaliy Filippov c0c01a8e57 Fix writeback detection in ublk server for --image mode 2026-06-10 00:31:48 +03:00
Vitaliy Filippov 334755e912 Fix vitastor-disk resize to smaller size clearing all data O_o, add a test for it 2026-06-10 00:30:42 +03:00
Vitaliy Filippov c16f955a51 Limit the amount of incoming data for NFS clients to prevent choking on memory 2026-06-06 13:56:20 +03:00
Vitaliy Filippov c1dc14f5ee Fix old store no_inode_stats not working 2026-06-05 00:50:53 +03:00
150 changed files with 6483 additions and 1856 deletions
+108
View File
@@ -306,6 +306,78 @@ jobs:
echo ""
done
test_dump_load:
runs-on: ubuntu-latest
needs: build
container: ${{env.TEST_IMAGE}}:${{github.sha}}
steps:
- name: Run test
id: test
timeout-minutes: 3
run: /root/vitastor/tests/test_dump_load.sh
- name: Print logs
if: always() && steps.test.outcome == 'failure'
run: |
for i in /root/vitastor/testdata/*.log /root/vitastor/testdata/*.txt; do
echo "-------- $i --------"
cat $i
echo ""
done
test_dump_load_32k:
runs-on: ubuntu-latest
needs: build
container: ${{env.TEST_IMAGE}}:${{github.sha}}
steps:
- name: Run test
id: test
timeout-minutes: 3
run: TEST_NAME=32k OSD_ARGS="--data_csum_type crc32c --csum_block_size 32k" OFFSET_ARGS="$OSD_ARGS" /root/vitastor/tests/test_dump_load.sh
- name: Print logs
if: always() && steps.test.outcome == 'failure'
run: |
for i in /root/vitastor/testdata/*.log /root/vitastor/testdata/*.txt; do
echo "-------- $i --------"
cat $i
echo ""
done
test_old_dump_load:
runs-on: ubuntu-latest
needs: build
container: ${{env.TEST_IMAGE}}:${{github.sha}}
steps:
- name: Run test
id: test
timeout-minutes: 3
run: OLD=1 /root/vitastor/tests/test_dump_load.sh
- name: Print logs
if: always() && steps.test.outcome == 'failure'
run: |
for i in /root/vitastor/testdata/*.log /root/vitastor/testdata/*.txt; do
echo "-------- $i --------"
cat $i
echo ""
done
test_dump_load_old_32k:
runs-on: ubuntu-latest
needs: build
container: ${{env.TEST_IMAGE}}:${{github.sha}}
steps:
- name: Run test
id: test
timeout-minutes: 3
run: TEST_NAME=old_32k OLD=1 OSD_ARGS="--data_csum_type crc32c --csum_block_size 32k" OFFSET_ARGS="$OSD_ARGS" /root/vitastor/tests/test_dump_load.sh
- name: Print logs
if: always() && steps.test.outcome == 'failure'
run: |
for i in /root/vitastor/testdata/*.log /root/vitastor/testdata/*.txt; do
echo "-------- $i --------"
cat $i
echo ""
done
test_old_interrupted_rebalance:
runs-on: ubuntu-latest
needs: build
@@ -1458,6 +1530,24 @@ jobs:
echo ""
done
test_resize_last:
runs-on: ubuntu-latest
needs: build
container: ${{env.TEST_IMAGE}}:${{github.sha}}
steps:
- name: Run test
id: test
timeout-minutes: 3
run: /root/vitastor/tests/test_resize_last.sh
- name: Print logs
if: always() && steps.test.outcome == 'failure'
run: |
for i in /root/vitastor/testdata/*.log /root/vitastor/testdata/*.txt; do
echo "-------- $i --------"
cat $i
echo ""
done
test_resize_auto:
runs-on: ubuntu-latest
needs: build
@@ -1494,6 +1584,24 @@ jobs:
echo ""
done
test_old_resize_last:
runs-on: ubuntu-latest
needs: build
container: ${{env.TEST_IMAGE}}:${{github.sha}}
steps:
- name: Run test
id: test
timeout-minutes: 3
run: OLD=1 /root/vitastor/tests/test_resize_last.sh
- name: Print logs
if: always() && steps.test.outcome == 'failure'
run: |
for i in /root/vitastor/testdata/*.log /root/vitastor/testdata/*.txt; do
echo "-------- $i --------"
cat $i
echo ""
done
test_old_resize_auto:
runs-on: ubuntu-latest
needs: build
+1 -1
View File
@@ -2,7 +2,7 @@ cmake_minimum_required(VERSION 2.8...3.30)
project(vitastor)
set(VITASTOR_VERSION "3.0.13")
set(VITASTOR_VERSION "3.0.15")
include(CTest)
+163 -92
View File
@@ -1,112 +1,183 @@
# Vitastor
# tromcho.net
[Читать на русском](README-ru.md)
Repository này chứa toàn bộ source code của website **tromcho.net**.
## The Idea
## Giới thiệu
Make Clustered Block Storage Fast Again.
`tromcho.net` là mã nguồn website được tổ chức để phục vụ phát triển, triển khai và vận hành theo quy trình chuẩn trên GitHub.
README này đóng vai trò tài liệu khởi đầu cho lập trình viên, DevOps engineer và cộng tác viên khi tiếp cận repository.
Vitastor is a distributed block, file and object SDS, direct replacement of Ceph RBD, CephFS and RGW,
and also internal SDS's of public clouds. However, in contrast to them, Vitastor is fast
and simple at the same time. The only thing is it's slightly young :-).
## Mục tiêu repository
Vitastor is architecturally similar to Ceph which means strong consistency,
primary-replication, symmetric clustering and automatic data distribution over any
number of drives of any size with configurable redundancy (replication or erasure codes/XOR).
- Quản lý tập trung toàn bộ source code của website.
- Chuẩn hóa quy trình phát triển, review và triển khai.
- Tạo nền tảng rõ ràng cho việc CI/CD, kiểm thử và vận hành production.
- Hỗ trợ onboarding nhanh cho thành viên mới.
Vitastor targets primarily SSD and SSD+HDD clusters with at least 10 Gbit/s network,
supports TCP and RDMA and may achieve 4 KB read and write latency as low as ~0.1 ms
with proper hardware which is ~10 times faster than other popular SDS's like Ceph
or internal systems of public clouds.
## Cấu trúc thư mục đề xuất
Vitastor supports QEMU, UBLK, NBD, NFS protocols, OpenStack, OpenNebula, Proxmox, Kubernetes drivers.
More drivers may be created easily.
```text
.
├── app/ # Source code ứng dụng chính
├── public/ # Static files, images, favicon, robots.txt
├── config/ # Cấu hình môi trường, app, service integration
├── database/ # Migration, seed, schema
├── tests/ # Unit test, integration test, e2e test
├── scripts/ # Script hỗ trợ build, deploy, backup, maintenance
├── docs/ # Tài liệu kỹ thuật, kiến trúc, quy trình
├── .github/ # GitHub Actions, issue template, PR template
├── Dockerfile # Build image ứng dụng
├── docker-compose.yml # Chạy local/dev bằng container
├── .env.example # Biến môi trường mẫu
└── README.md
```
Read more details in the documentation. You can start from here: [Quick Start](docs/intro/quickstart.en.md).
> Cấu trúc thực tế có thể thay đổi theo framework đang sử dụng.
## Talks and presentations
## Yêu cầu môi trường
- KuberConf'2025: [video](https://vitastor.io/presentation/kuberconf.webm)
- Highload'2025: [video](https://vitastor.io/presentation/hl2025/hl2025.webm),
[youtube](https://www.youtube.com/watch?v=0R8MLjFtz7g), presentation
([in Russian](https://vitastor.io/presentation/hl2025/), [in English](https://vitastor.io/presentation/hl2025/en.html))
- Highload'2022: presentation ([in Russian](https://vitastor.io/presentation/highload/highload.html)),
[video](https://vitastor.io/presentation/highload/talk.webm)
- DevOpsConf'2021: presentation ([in Russian](https://vitastor.io/presentation/devopsconf/devopsconf.html),
[in English](https://vitastor.io/presentation/devopsconf/devopsconf_en.html)),
[video](https://vitastor.io/presentation/devopsconf/talk.webm)
Tùy theo stack công nghệ của website, môi trường phát triển nên có:
## Documentation
- Git
- Docker và Docker Compose
- Node.js / PHP / Python / runtime phù hợp với dự án
- Make (khuyến nghị)
- Truy cập vào file cấu hình môi trường `.env`
- Introduction
- [Quick Start](docs/intro/quickstart.en.md)
- [Features](docs/intro/features.en.md)
- [Architecture](docs/intro/architecture.en.md)
- [Author and license](docs/intro/author.en.md)
- Installation
- [Packages](docs/installation/packages.en.md)
- [Docker](docs/installation/docker.en.md)
- [Proxmox](docs/installation/proxmox.en.md)
- [OpenNebula](docs/installation/opennebula.en.md)
- [OpenStack](docs/installation/openstack.en.md)
- [Kubernetes CSI](docs/installation/kubernetes.en.md)
- [S3](docs/installation/s3.en.md)
- [Building from Source](docs/installation/source.en.md)
- Configuration
- [Overview](docs/config.en.md)
- Parameter Reference
- [Common](docs/config/common.en.md)
- [Network](docs/config/network.en.md)
- [Client](docs/config/client.en.md)
- [Global Disk Layout](docs/config/layout-cluster.en.md)
- [OSD Disk Layout](docs/config/layout-osd.en.md)
- [OSD Runtime Parameters](docs/config/osd.en.md)
- [Monitor](docs/config/monitor.en.md)
- [Pool configuration](docs/config/pool.en.md)
- [Image metadata in etcd](docs/config/inode.en.md)
- Usage
- [vitastor-cli](docs/usage/cli.en.md) (command-line interface)
- [vitastor-disk](docs/usage/disk.en.md) (disk management tool)
- [fio](docs/usage/fio.en.md) for benchmarks
- [UBLK](docs/usage/ublk.en.md) for kernel mounts
- [NBD](docs/usage/nbd.en.md) - old interface for kernel mounts
- [QEMU, qemu-img and VDUSE](docs/usage/qemu.en.md)
- [NFS](docs/usage/nfs.en.md) clustered file system and pseudo-FS proxy
- [Administration](docs/usage/admin.en.md)
- Performance
- [Understanding storage performance](docs/performance/understanding.en.md)
- [Theoretical performance](docs/performance/theoretical.en.md)
- [Example comparison with Ceph](docs/performance/comparison1.en.md)
- [Newer benchmark of Vitastor 1.3.1](docs/performance/bench2.en.md)
## Bắt đầu nhanh
## Author and License
### 1. Clone repository
Copyright (c) Vitaliy Filippov (vitalif [at] yourcmc.ru), 2019+
```bash
git clone https://github.com/<your-org>/tromcho.net.git
cd tromcho.net
```
Join Vitastor Telegram Chat: https://t.me/vitastor
### 2. Tạo file môi trường
All server-side code (OSD, Monitor and so on) is licensed under the terms of
Vitastor Network Public License 1.1 (VNPL 1.1), a copyleft license based on
GNU GPLv3.0 with the additional "Network Interaction" clause which requires
opensourcing all programs directly or indirectly interacting with Vitastor
through a computer network and expressly designed to be used in conjunction
with it ("Proxy Programs"). Proxy Programs may be made public not only under
the terms of the same license, but also under the terms of any GPL-Compatible
Free Software License, as listed by the Free Software Foundation.
This is a stricter copyleft license than the Affero GPL.
```bash
cp .env.example .env
```
Please note that VNPL doesn't require you to open the code of proprietary
software running inside a VM if it's not specially designed to be used with
Vitastor.
Sau đó cập nhật các biến cấu hình cần thiết trong file `.env`.
Basically, you can't use the software in a proprietary environment to provide
its functionality to users without opensourcing all intermediary components
standing between the user and Vitastor or purchasing a commercial license
from the author 😀.
### 3. Chạy môi trường local
Client libraries (cluster_client and so on) are dual-licensed under the same
VNPL 1.1 and also GNU GPL 2.0 or later to allow for compatibility with GPLed
software like QEMU and fio.
Nếu dự án dùng Docker:
You can find the full text of VNPL-1.1 in the file [VNPL-1.1.txt](VNPL-1.1.txt).
GPL 2.0 is also included in this repository as [GPL-2.0.txt](GPL-2.0.txt).
```bash
docker compose up -d --build
```
Nếu dự án chạy trực tiếp theo framework, sử dụng lệnh tương ứng của stack hiện tại.
## Quy trình phát triển
- Tạo branch mới từ `main` hoặc `develop`.
- Đặt tên branch rõ ràng, ví dụ: `feature/homepage-banner`, `fix/login-timeout`.
- Commit ngắn gọn, đúng ngữ cảnh thay đổi.
- Tạo Pull Request để review trước khi merge.
- Không commit file bí mật như `.env`, private key hoặc credential.
## Quy ước commit
Khuyến nghị dùng convention sau:
```text
feat: thêm chức năng mới
fix: sửa lỗi
refactor: tái cấu trúc mã nguồn
chore: cập nhật tác vụ phụ trợ
ci: thay đổi pipeline CI/CD
docs: cập nhật tài liệu
test: bổ sung hoặc cập nhật kiểm thử
```
## CI/CD
Repository nên tích hợp các bước tự động sau:
- Lint source code
- Chạy unit test / integration test
- Build artifact hoặc Docker image
- Scan bảo mật dependency/container
- Deploy tới staging hoặc production theo rule xác định
Ví dụ vị trí cấu hình pipeline:
```text
.github/workflows/
```
## Biến môi trường
Không commit file `.env` thật lên GitHub.
Nên cung cấp `.env.example` với:
- Danh sách biến bắt buộc
- Giá trị mẫu an toàn
- Ghi chú ngắn cho từng biến quan trọng
Ví dụ:
```env
APP_ENV=local
APP_DEBUG=true
APP_URL=http://localhost
DB_HOST=127.0.0.1
DB_PORT=3306
DB_NAME=tromcho
DB_USER=user
DB_PASSWORD=change_me
```
## Triển khai
Khuyến nghị tách rõ các môi trường:
- local
- development
- staging
- production
Các thành phần nên được chuẩn hóa khi triển khai:
- Biến môi trường
- Reverse proxy / web server
- TLS certificate
- Database migration
- Backup strategy
- Log rotation và monitoring
## Bảo mật
- Không đưa secrets vào source code.
- Bật branch protection cho nhánh quan trọng.
- Review dependency định kỳ.
- Áp dụng nguyên tắc least privilege cho tài khoản deploy.
- Theo dõi log, audit và cảnh báo bất thường.
## Đóng góp
Khi đóng góp vào repository:
1. Fork hoặc tạo branch làm việc.
2. Cập nhật mã nguồn theo phạm vi thay đổi.
3. Kiểm tra local trước khi tạo Pull Request.
4. Viết mô tả PR rõ ràng: mục tiêu, phạm vi ảnh hưởng, cách kiểm thử.
## Tài liệu nên bổ sung
Repository này nên có thêm các tài liệu sau trong thư mục `docs/`:
- Kiến trúc hệ thống
- Sơ đồ database
- Luồng deploy
- Quy trình backup/restore
- Hướng dẫn xử lý sự cố
- Checklist release
## License
Copyright 2026 Trộm Chó chấm Nét
---
+1 -1
View File
@@ -1,4 +1,4 @@
VITASTOR_VERSION ?= v3.0.13
VITASTOR_VERSION ?= v3.0.15
all: build push
+1 -1
View File
@@ -49,7 +49,7 @@ spec:
capabilities:
add: ["SYS_ADMIN"]
allowPrivilegeEscalation: true
image: vitalif/vitastor-csi:v3.0.13
image: vitalif/vitastor-csi:v3.0.15
args:
- "--node=$(NODE_ID)"
- "--endpoint=$(CSI_ENDPOINT)"
+1 -1
View File
@@ -121,7 +121,7 @@ spec:
privileged: true
capabilities:
add: ["SYS_ADMIN"]
image: vitalif/vitastor-csi:v3.0.13
image: vitalif/vitastor-csi:v3.0.15
args:
- "--node=$(NODE_ID)"
- "--endpoint=$(CSI_ENDPOINT)"
+1 -1
View File
@@ -5,7 +5,7 @@ package vitastor
const (
vitastorCSIDriverName = "csi.vitastor.io"
vitastorCSIDriverVersion = "3.0.13"
vitastorCSIDriverVersion = "3.0.15"
)
// Config struct fills the parameters of request or user input
+1 -1
View File
@@ -1,4 +1,4 @@
vitastor (3.0.13-1) unstable; urgency=medium
vitastor (3.0.15-1) unstable; urgency=medium
* Bugfixes
+1
View File
@@ -11,6 +11,7 @@ override_dh_install:
cp -v node-binding/package.json node-binding/index.js node-binding/addon.cc node-binding/addon.h node-binding/client.cc node-binding/client.h debian/tmp/usr/lib/x86_64-linux-gnu/nodejs/vitastor
cp -v node-binding/build/Release/addon.node debian/tmp/usr/lib/x86_64-linux-gnu/nodejs/vitastor/build/Release
dh_install
cd debian/vitastor-mon/usr/lib/vitastor/mon && npm install --production
override_dh_installdeb:
cat debian/fio_version >> debian/vitastor-fio.substvars
-6
View File
@@ -37,12 +37,6 @@ rm -rf a b
echo "dep:fio=$FIO" > debian/fio_version
cd /root/vitastor/packages/vitastor-$REL/vitastor-$VER
mkdir mon/node_modules
cd mon/node_modules
curl -s https://git.yourcmc.ru/vitalif/antietcd/archive/master.tar.gz | tar -zx
curl -s https://git.yourcmc.ru/vitalif/tinyraft/archive/master.tar.gz | tar -zx
cd /root/vitastor/packages/vitastor-$REL
if [[ ( "$REL" = "trixie" || "$REL" = "resolute" ) && -e ../vitastor-bookworm/vitastor_$VER.orig.tar.xz ]]; then
# Fucking shit, archives differ between bookworm (xz 5.4.1) and trixie (xz 5.8.1)
+1 -1
View File
@@ -1,4 +1,4 @@
VITASTOR_VERSION ?= v3.0.13
VITASTOR_VERSION ?= v3.0.15
all: build push
+1 -1
View File
@@ -4,7 +4,7 @@
#
# Desired Vitastor version
VITASTOR_VERSION=v3.0.13
VITASTOR_VERSION=v3.0.15
# Additional arguments for all containers
# For example, you may want to specify a custom logging driver here
+2 -2
View File
@@ -26,9 +26,9 @@ at Vitastor Kubernetes operator: https://github.com/Antilles7227/vitastor-operat
The instruction is very simple.
1. Download a Docker image of the desired version: \
`docker pull vitalif/vitastor:v3.0.13`
`docker pull vitalif/vitastor:v3.0.15`
2. Install scripts to the host system: \
`docker run --rm -it -v /etc:/host-etc -v /usr/bin:/host-bin vitalif/vitastor:v3.0.13 install.sh`
`docker run --rm -it -v /etc:/host-etc -v /usr/bin:/host-bin vitalif/vitastor:v3.0.15 install.sh`
3. Reload udev rules: \
`udevadm control --reload-rules`
4. Enable the vitastor-host service: \
+2 -2
View File
@@ -25,9 +25,9 @@ Vitastor можно установить в Docker/Podman. При этом etcd,
Инструкция по установке максимально простая.
1. Скачайте Docker-образ желаемой версии: \
`docker pull vitalif/vitastor:v3.0.13`
`docker pull vitalif/vitastor:v3.0.15`
2. Установите скрипты в хост-систему командой: \
`docker run --rm -it -v /etc:/host-etc -v /usr/bin:/host-bin vitalif/vitastor:v3.0.13 install.sh`
`docker run --rm -it -v /etc:/host-etc -v /usr/bin:/host-bin vitalif/vitastor:v3.0.15 install.sh`
3. Перезагрузите правила udev: \
`udevadm control --reload-rules`
4. Включите сервис vitastor-host: \
+19 -7
View File
@@ -18,7 +18,7 @@ class AntiEtcdAdapter
cluster = cluster ? (''+(cluster||'')).split(/,+/) : [];
cluster = Object.keys(cluster.reduce((a, url) =>
{
a[url.toLowerCase().replace(/^(https?:\/\/)/, '').replace(/\/.*$/, '')] = true;
a[url.toLowerCase().replace(/^(https?:\/\/)?(.*?)(\/.*)?$/, (m, m1, m2) => (m1||'http://')+m2)] = true;
return a;
}, {}));
const cfg_port = config.antietcd_port;
@@ -26,7 +26,18 @@ class AntiEtcdAdapter
is_local['0.0.0.0'] = true;
is_local['::'] = true;
is_local[''] = true;
const selected = cluster.map(s => s.split(':', 2)).filter(ip => is_local[ip[0]] && (!cfg_port || ip[1] == cfg_port));
// split :, 3 -> <schema>:<//ip>:<port>
const selected = [];
for (let i = 0; i < cluster.length; i++)
{
const m = /^(https?:\/\/)?(?:\[(.*)\]|([^\[\:]+))(?::(\d+))?$/.exec(cluster[i]);
if (!m)
continue;
const ip = m[3] || m[2];
const port = m[4] || 2379;
if (is_local[ip] && (!cfg_port || port == cfg_port))
selected.push({ idx: i, ip, port });
}
if (selected.length > 1)
{
console.error('More than 1 etcd_address matches local IPs, please specify port');
@@ -35,15 +46,16 @@ class AntiEtcdAdapter
else if (selected.length == 1)
{
const antietcd_config = {
ip: selected[0][0],
port: selected[0][1],
data: config.antietcd_data_file || ((config.antietcd_data_dir || '/var/lib/vitastor') + '/mon_'+selected[0][1]+'.json.gz'),
ip: selected[0].ip,
port: selected[0].port,
data: config.antietcd_data_file || ((config.antietcd_data_dir || '/var/lib/vitastor') + '/mon_'+selected[0].port+'.json.gz'),
persist_filter: vitastor_persist_filter({ vitastor_prefix: config.etcd_prefix || '/vitastor' }),
node_id: selected[0][0]+':'+selected[0][1], // node_id = ip:port
cluster: (cluster.length == 1 ? null : cluster.reduce((a, c) => { a[c] = "http://"+c; return a; }, {})),
node_id: cluster[selected[0].idx].replace(/^(https?:\/\/)/, ''), // same as in <cluster> below
cluster: (cluster.length == 1 ? null : cluster.reduce((a, c) => { a[c.replace(/^(https?:\/\/)/, '')] = c; return a; }, {})),
cluster_key: (config.etcd_prefix || '/vitastor'),
stale_read: 1,
log_level: 1,
logs: { cluster: true },
};
for (const key in config)
{
+2 -2
View File
@@ -1,6 +1,6 @@
{
"name": "vitastor-mon",
"version": "3.0.13",
"version": "3.0.15",
"description": "Vitastor SDS monitor service",
"main": "mon-main.js",
"scripts": {
@@ -9,7 +9,7 @@
"author": "Vitaliy Filippov",
"license": "UNLICENSED",
"dependencies": {
"antietcd": "^1.2.4",
"antietcd": "^1.3.1",
"sprintf-js": "^1.1.2",
"ws": "^7.2.5"
},
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "vitastor",
"version": "3.0.13",
"version": "3.0.15",
"description": "Low-level native bindings to Vitastor client library",
"main": "index.js",
"keywords": [
+45 -10
View File
@@ -366,15 +366,38 @@ sub map_volume
my $prefix = defined $scfg->{vitastor_prefix} ? $scfg->{vitastor_prefix} : 'pve/';
my ($vtype, $img_name, $vmid) = $class->parse_volname($volname);
my $name = $img_name;
my $name = $prefix.$img_name;
$name .= '@'.$snapname if $snapname;
my $mapped = run_cli($scfg, [ 'ls' ], binary => '/usr/bin/vitastor-nbd');
my ($kerneldev) = grep { $mapped->{$_}->{image} eq $prefix.$name } keys %$mapped;
return $kerneldev if $kerneldev && -b $kerneldev; # already mapped
my ($kerneldev) = grep {
$mapped->{$_} && $mapped->{$_}->{image} && $mapped->{$_}->{image} eq $name
} keys %$mapped;
$kerneldev = run_cli($scfg, [ 'map', '--image', $prefix.$name ], binary => '/usr/bin/vitastor-nbd', json => 0);
return $kerneldev;
if ($kerneldev && -b $kerneldev)
{
my $size = `/usr/sbin/blockdev --getsize64 $kerneldev`;
return $kerneldev if $size && $size > 0;
}
my $map_out = run_cli($scfg, [ 'map', '--image', $name ], binary => '/usr/bin/vitastor-nbd', json => 0);
$map_out =~ s/^\s+|\s+$//gso;
# Wait until the device is started
for (my $i = 0; $i < 100; $i++)
{
$mapped = run_cli($scfg, [ 'ls' ], binary => '/usr/bin/vitastor-nbd');
($kerneldev) = grep { $mapped->{$_} && $mapped->{$_}->{image} && $mapped->{$_}->{image} eq $name } keys %$mapped;
if ($kerneldev && -b $kerneldev)
{
my $size = `/usr/sbin/blockdev --getsize64 $kerneldev`;
return $kerneldev if $size && $size > 0;
}
select(undef, undef, undef, 0.1);
}
die "Failed to map Vitastor image $name via NBD".
($map_out ? ", vitastor-nbd map returned '$map_out'" : "")."\n";
}
sub unmap_volume
@@ -383,13 +406,19 @@ sub unmap_volume
my $prefix = defined $scfg->{vitastor_prefix} ? $scfg->{vitastor_prefix} : 'pve/';
my ($vtype, $name, $vmid) = $class->parse_volname($volname);
$name = $prefix.$name;
$name .= '@'.$snapname if $snapname;
my $mapped = run_cli($scfg, [ 'ls' ], binary => '/usr/bin/vitastor-nbd');
my ($kerneldev) = grep { $mapped->{$_}->{image} eq $prefix.$name } keys %$mapped;
if ($kerneldev && -b $kerneldev)
my @kerneldevs = grep {
$mapped->{$_} && $mapped->{$_}->{image} && $mapped->{$_}->{image} eq $name
} keys %$mapped;
for my $kerneldev (@kerneldevs)
{
run_cli($scfg, [ 'unmap', $kerneldev ], binary => '/usr/bin/vitastor-nbd', json => 0);
next if !$kerneldev || !-b $kerneldev;
eval { run_cli($scfg, [ 'unmap', $kerneldev ], binary => '/usr/bin/vitastor-nbd', json => 0); };
warn "Failed to unmap Vitastor image $name from $kerneldev: $@" if $@;
}
return 1;
@@ -405,7 +434,13 @@ sub activate_volume
sub deactivate_volume
{
my ($class, $storeid, $scfg, $volname, $snapname, $cache) = @_;
$class->unmap_volume($storeid, $scfg, $volname, $snapname) if $scfg->{vitastor_nbd};
# Even with vitastor_nbd=0, Proxmox may call map_volume() for special
# volumes like tpmstate0 because swtpm needs a local file/block path.
# Therefore, always try to unmap an existing NBD mapping here.
# unmap_volume() is a no-op if the volume is not currently mapped.
$class->unmap_volume($storeid, $scfg, $volname, $snapname);
return 1;
}
+1 -1
View File
@@ -50,7 +50,7 @@ from cinder.volume import configuration
from cinder.volume import driver
from cinder.volume import volume_utils
VITASTOR_VERSION = '3.0.13'
VITASTOR_VERSION = '3.0.15'
LOG = logging.getLogger(__name__)
+2 -2
View File
@@ -1,11 +1,11 @@
Name: vitastor
Version: 3.0.13
Version: 3.0.15
Release: 1%{?dist}
Summary: Vitastor, a fast software-defined clustered block storage
License: Vitastor Network Public License 1.1
URL: https://vitastor.io/
Source0: vitastor-3.0.13.el10.tar.gz
Source0: vitastor-3.0.15.el10.tar.gz
BuildRequires: gperftools-devel
BuildRequires: gcc-c++
+2 -2
View File
@@ -1,11 +1,11 @@
Name: vitastor
Version: 3.0.13
Version: 3.0.15
Release: 1%{?dist}
Summary: Vitastor, a fast software-defined clustered block storage
License: Vitastor Network Public License 1.1
URL: https://vitastor.io/
Source0: vitastor-3.0.13.el7.tar.gz
Source0: vitastor-3.0.15.el7.tar.gz
BuildRequires: gperftools-devel
BuildRequires: devtoolset-9-gcc-c++
+2 -2
View File
@@ -1,11 +1,11 @@
Name: vitastor
Version: 3.0.13
Version: 3.0.15
Release: 1%{?dist}
Summary: Vitastor, a fast software-defined clustered block storage
License: Vitastor Network Public License 1.1
URL: https://vitastor.io/
Source0: vitastor-3.0.13.el8.tar.gz
Source0: vitastor-3.0.15.el8.tar.gz
BuildRequires: gperftools-devel
BuildRequires: gcc-toolset-9-gcc-c++
+2 -2
View File
@@ -1,11 +1,11 @@
Name: vitastor
Version: 3.0.13
Version: 3.0.15
Release: 1%{?dist}
Summary: Vitastor, a fast software-defined clustered block storage
License: Vitastor Network Public License 1.1
URL: https://vitastor.io/
Source0: vitastor-3.0.13.el9.tar.gz
Source0: vitastor-3.0.15.el9.tar.gz
BuildRequires: gperftools-devel
BuildRequires: gcc-c++
+1 -1
View File
@@ -20,7 +20,7 @@ if("${CMAKE_INSTALL_PREFIX}" MATCHES "^/usr/local/?$")
endif()
set(ENABLE_COVERAGE false CACHE BOOL "Enable code coverage")
add_definitions(-DVITASTOR_VERSION="3.0.13")
add_definitions(-DVITASTOR_VERSION="3.0.15")
add_definitions(-D_GNU_SOURCE -D_LARGEFILE64_SOURCE -D_FILE_OFFSET_BITS=64 -Wall -Wno-sign-compare -Wno-comment -Wno-parentheses -Wno-pointer-arith -fdiagnostics-color=always -fno-omit-frame-pointer -fvisibility=hidden -I ${CMAKE_SOURCE_DIR}/src)
add_link_options(-fno-omit-frame-pointer)
if (${WITH_ASAN})
+12 -8
View File
@@ -541,17 +541,21 @@ int blockstore_disk_t::trim_data(std::function<bool(uint64_t)> is_used)
if (range[0] % discard_granularity)
range[0] = range[0] + discard_granularity - (range[0] % discard_granularity);
if (range[0] >= range[1])
continue;
range[1] -= range[0];
range[1] = 0;
else
range[1] -= range[0];
}
r = ioctl(data_fd, BLKDISCARD, &range);
if (r != 0)
if (range[1] > 0)
{
fprintf(stderr, "Failed to execute BLKDISCARD %ju+%ju on %s: %s (code %d)\n",
range[0], range[1], data_device.c_str(), strerror(-r), r);
return -errno;
r = ioctl(data_fd, BLKDISCARD, &range);
if (r != 0)
{
fprintf(stderr, "Failed to execute BLKDISCARD %ju+%ju on %s: %s (code %d)\n",
range[0], range[1], data_device.c_str(), strerror(-r), r);
return -errno;
}
discarded += range[1];
}
discarded += range[1];
}
j = i+1;
}
+50 -23
View File
@@ -22,7 +22,7 @@
#define HEAP_INFLIGHT_DONE 1
#define HEAP_INFLIGHT_COMPACTABLE 2
#define HEAP_INFLIGHT_COMPACTED 4
#define HEAP_INFLIGHT_OVERWRITE 4
#define HEAP_INFLIGHT_GC 8
#define HEAP_INFLIGHT_EXPLICIT 16
@@ -122,6 +122,8 @@ uint32_t heap_entry_t::get_size(blockstore_heap_t *heap)
}
if (type() == BS_HEAP_SMALL_WRITE || type() == BS_HEAP_INTENT_WRITE)
{
if (size < sizeof(heap_small_write_t))
return heap->get_small_entry_size(0, 0);
return heap->get_small_entry_size(small().offset, small().len);
}
return heap->get_simple_entry_size();
@@ -364,14 +366,11 @@ corrupted_object:
return EDOM;
}
}
if (((wr->entry_type & BS_HEAP_TYPE) == BS_HEAP_SMALL_WRITE ||
(wr->entry_type & BS_HEAP_TYPE) == BS_HEAP_INTENT_WRITE) &&
wr->size < sizeof(heap_small_write_t))
if (wr->size != wr->get_size(this))
{
// Small writes require accessing offset & len to calculate correct length,
// so require at least sizeof(heap_small_write_t) for them
fprintf(stderr, "Error: entry %jx:%jx v%ju has invalid size in metadata block %u at %u (%u < min %zu bytes)\n",
wr->inode, wr->stripe, wr->version, block_num, block_offset, wr->size, sizeof(heap_small_write_t));
// Check entry size
fprintf(stderr, "Error: entry %jx:%jx v%ju has invalid size in metadata block %u at %u (%u != %u bytes)\n",
wr->inode, wr->stripe, wr->version, block_num, block_offset, wr->size, wr->get_size(this));
goto corrupted_object;
}
if (wr->entry_type == BS_HEAP_COMMIT && !wr->version)
@@ -683,10 +682,14 @@ int blockstore_heap_t::mark_used_blocks()
}
use_data(wr->inode, wr->big_location(this));
}
if (wr->is_compactable() && !added)
if (wr->is_compactable())
{
compact_queue.push_back((object_id){ .inode = wr->inode, .stripe = wr->stripe });
added = true;
to_compact_count++;
if (!added)
{
compact_queue.push_back((object_id){ .inode = wr->inode, .stripe = wr->stripe });
added = true;
}
}
if (wr->is_overwrite())
{
@@ -696,6 +699,11 @@ int blockstore_heap_t::mark_used_blocks()
});
}
}
for (auto li: init_erase_items)
{
unlink_list_item(li);
}
init_erase_items.clear();
if (dsk->gc_on_start)
{
recheck_full_gc();
@@ -731,7 +739,6 @@ void blockstore_heap_t::init_erase_bad_entry(heap_list_item_t *li)
inf.garbage_space -= (li->entry.is_garbage() ? li->entry.size : 0);
});
recheck_modified_blocks.insert(li->block_num);
unlink_list_item(li);
}
bool blockstore_heap_t::init_erase_double_claim(heap_list_item_t *prev_li, heap_list_item_t *cur_li)
@@ -786,6 +793,8 @@ bool blockstore_heap_t::init_erase_double_claim(heap_list_item_t *prev_li, heap_
overwritten = erase_li->entry.is_overwrite();
}
init_erase_bad_entry(erase_li);
// Can't erase (mutate map) while iterating, so postpone it
init_erase_items.push_back(erase_li);
erase_li = prev_erase_li;
}
}
@@ -799,6 +808,8 @@ bool blockstore_heap_t::init_erase_double_claim(heap_list_item_t *prev_li, heap_
auto next_erase_li = erase_li->next;
init_free_bad_entry(&erase_li->entry);
init_erase_bad_entry(erase_li);
// Can't erase (mutate map) while iterating, so postpone it
init_erase_items.push_back(erase_li);
erase_li = next_erase_li;
}
erase_li = cur_li;
@@ -807,6 +818,8 @@ bool blockstore_heap_t::init_erase_double_claim(heap_list_item_t *prev_li, heap_
{
auto prev_erase_li = erase_li->prev;
init_erase_bad_entry(erase_li);
// Can't erase (mutate map) while iterating, so postpone it
init_erase_items.push_back(erase_li);
erase_li = prev_erase_li;
}
}
@@ -877,6 +890,7 @@ void blockstore_heap_t::recheck_drop_entries(heap_entry_t *obj, heap_entry_t *ba
auto prev = li->prev;
assert(li->entry.type() == bad_wr->type());
init_erase_bad_entry(li);
unlink_list_item(li);
li = prev;
}
}
@@ -1139,7 +1153,7 @@ bool blockstore_heap_t::calc_block_checksums(uint32_t *block_csums, uint8_t *bit
while (pos < end && pos < block_end && !(bitmap[pos/dsk->bitmap_granularity/8] & (1 << ((pos/dsk->bitmap_granularity) % 8))))
pos += dsk->bitmap_granularity;
// zero padding at the beginning or at the end of the block is not counted
if (pos > prev && prev > 0 && pos < block_end)
if (pos > prev && prev > blk_start && pos < block_end)
block_crc = crc32c_pad(block_crc, NULL, 0, pos-prev, 0);
prev = pos;
while (pos < end && pos < block_end && (bitmap[pos/dsk->bitmap_granularity/8] & (1 << ((pos/dsk->bitmap_granularity) % 8))))
@@ -1537,7 +1551,7 @@ int blockstore_heap_t::add_entry(uint32_t wr_size, uint32_t *modified_block,
// Remember the object as dirty and remove older entries when this block is written and fsynced
push_inflight_lsn(next_lsn, new_wr,
(explicit_complete ? HEAP_INFLIGHT_EXPLICIT : 0) |
(new_wr->is_overwrite() ? HEAP_INFLIGHT_COMPACTED : 0) |
(new_wr->is_overwrite() ? HEAP_INFLIGHT_OVERWRITE : 0) |
(new_wr->is_compactable() ? HEAP_INFLIGHT_COMPACTABLE : 0));
insert_list_items(&li, 1, false);
li->block_num = block_num;
@@ -1571,10 +1585,22 @@ int blockstore_heap_t::add_small_write(object_id oid, heap_entry_t **obj_ptr, ui
wr->small().location = location;
if (bitmap)
memcpy(wr->get_ext_bitmap(this), bitmap, dsk->clean_entry_bitmap_size);
else if (obj)
memcpy(wr->get_ext_bitmap(this), obj->get_ext_bitmap(this), dsk->clean_entry_bitmap_size);
else
memset(wr->get_ext_bitmap(this), 0, dsk->clean_entry_bitmap_size);
{
bool found = false;
iterate_with_stable(obj, UINT64_MAX, [&](heap_entry_t *old_wr, bool stable)
{
if (old_wr->get_ext_bitmap(this))
{
found = true;
memcpy(wr->get_ext_bitmap(this), old_wr->get_ext_bitmap(this), dsk->clean_entry_bitmap_size);
return false;
}
return true;
});
if (!found)
memset(wr->get_ext_bitmap(this), 0, dsk->clean_entry_bitmap_size);
}
calc_checksums(wr, (uint8_t*)data, true);
*obj_ptr = wr;
});
@@ -2274,7 +2300,7 @@ void blockstore_heap_t::use_data(inode_t inode, uint64_t location)
{
auto sh_it = pool_shard_settings.find(INODE_POOL(inode));
if (sh_it != pool_shard_settings.end() && sh_it->second.no_inode_stats)
inode = (INODE_POOL(inode) << POOL_ID_BITS);
inode = INODE_WITH_POOL(INODE_POOL(inode), 0);
assert(!data_alloc->get(location / dsk->data_block_size));
data_alloc->set(location / dsk->data_block_size, true);
inode_space_stats[inode] += dsk->data_block_size;
@@ -2285,7 +2311,7 @@ void blockstore_heap_t::free_data(inode_t inode, uint64_t location)
{
auto sh_it = pool_shard_settings.find(INODE_POOL(inode));
if (sh_it != pool_shard_settings.end() && sh_it->second.no_inode_stats)
inode = (INODE_POOL(inode) << POOL_ID_BITS);
inode = INODE_WITH_POOL(INODE_POOL(inode), 0);
assert(data_alloc->get(location / dsk->data_block_size));
data_alloc->set(location / dsk->data_block_size, false);
auto sp_it = inode_space_stats.find(inode);
@@ -2322,7 +2348,8 @@ void blockstore_heap_t::use_buffer_area(inode_t inode, uint64_t location, uint64
return;
}
assert(!(size % dsk->bitmap_granularity));
buffer_alloc->use(location / dsk->bitmap_granularity, size / dsk->bitmap_granularity);
bool ok = buffer_alloc->use(location / dsk->bitmap_granularity, size / dsk->bitmap_granularity);
assert(ok);
buffer_area_used_space += size;
}
@@ -2364,7 +2391,7 @@ void blockstore_heap_t::get_meta_block(uint32_t block_num, uint8_t *buffer)
}
}
void blockstore_heap_t::fill_block_empty_space(uint8_t *buffer, uint32_t pos)
void blockstore_heap_t::fill_block_empty_space(uint8_t *buffer, uint64_t pos)
{
if (pos > dsk->meta_block_size)
{
@@ -2452,7 +2479,7 @@ uint64_t blockstore_heap_t::get_garbage_memory()
void blockstore_heap_t::push_inflight_lsn(uint64_t lsn, heap_entry_t *wr, uint64_t flags)
{
uint64_t next_inf = first_inflight_lsn + inflight_lsn.size();
if (flags & (HEAP_INFLIGHT_COMPACTABLE|HEAP_INFLIGHT_COMPACTED))
if (flags & (HEAP_INFLIGHT_COMPACTABLE|HEAP_INFLIGHT_OVERWRITE))
{
to_compact_count++;
}
@@ -2519,7 +2546,7 @@ void blockstore_heap_t::mark_lsn_fsynced(uint64_t lsn)
void blockstore_heap_t::apply_inflight(heap_inflight_lsn_t & inflight)
{
auto wr = inflight.wr;
if (inflight.flags & HEAP_INFLIGHT_COMPACTED)
if (inflight.flags & HEAP_INFLIGHT_OVERWRITE)
{
// Mark previous entries as garbage, sequentially
mark_garbage_up_to(wr);
+2 -1
View File
@@ -219,6 +219,7 @@ class blockstore_heap_t
bool marked_used_blocks = false;
bool recheck_queue_filled = false;
std::vector<heap_list_item_t*> postponed_items;
std::vector<heap_list_item_t*> init_erase_items;
std::set<uint32_t> recheck_modified_blocks;
std::deque<heap_entry_t*> recheck_queue;
std::map<heap_entry_t*, heap_recheck_state_t> recheck_states;
@@ -361,7 +362,7 @@ public:
// get metadata block data buffer and used space
void get_meta_block(uint32_t block_num, uint8_t *buffer);
void fill_block_empty_space(uint8_t *buffer, uint32_t pos);
void fill_block_empty_space(uint8_t *buffer, uint64_t pos);
uint32_t get_meta_block_used_space(uint32_t block_num);
// get space usage statistics
+5 -2
View File
@@ -117,9 +117,12 @@ public:
journal_flusher_t *flusher;
int write_iodepth = 0;
int inflight_big = 0;
int intent_write_counter = 0;
bool fsyncing_data = false;
uint64_t data_fsync_next = 0;
uint64_t data_fsync_cur = 0;
uint64_t data_fsync_sent = 0;
uint64_t data_fsync_done = 0;
std::deque<bool> data_fsyncs;
bool live = false, queue_stall = false;
ring_loop_i *ringloop = NULL;
+61 -52
View File
@@ -10,7 +10,6 @@
#define INIT_META_EMPTY 0
#define INIT_META_READING 1
#define INIT_META_READ_DONE 2
#define INIT_META_WRITING 3
#define GET_SQE() \
sqe = bs->get_sqe();\
@@ -23,14 +22,15 @@ blockstore_init_meta::blockstore_init_meta(blockstore_impl_t *bs)
this->bs = bs;
}
void blockstore_init_meta::handle_event(ring_data_t *data, int buf_num)
void blockstore_init_meta::handle_event(ring_data_t *data, int buf_num, const char *op)
{
if (data->res < 0)
if (data->res != data->iov.iov_len)
{
throw std::runtime_error(
std::string("read metadata failed at offset ") + std::to_string(buf_num >= 0 ? bufs[buf_num].offset : last_read_offset) +
std::string(": ") + strerror(-data->res)
);
throw std::runtime_error(strprintf(
"%s failed at offset %ju: got %s (code %d), but expected %zu",
op, (buf_num >= 0 ? bufs[buf_num].offset : last_read_offset), strerror(-data->res),
data->res, data->iov.iov_len
));
}
if (buf_num >= 0)
{
@@ -60,7 +60,7 @@ int blockstore_init_meta::loop()
GET_SQE();
last_read_offset = 0;
data->iov = { bs->meta_superblock, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "read metadata header"); };
io_uring_prep_readv(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset);
bs->ringloop->submit();
submitted++;
@@ -72,25 +72,19 @@ resume_1:
}
if (is_zero((uint64_t*)bs->meta_superblock, bs->dsk.meta_block_size))
{
{
blockstore_meta_header_v3_t *hdr = (blockstore_meta_header_v3_t *)bs->meta_superblock;
hdr->zero = 0;
hdr->magic = BLOCKSTORE_META_MAGIC_V1;
hdr->version = bs->dsk.meta_format;
hdr->meta_block_size = bs->dsk.meta_block_size;
hdr->data_block_size = bs->dsk.data_block_size;
hdr->bitmap_granularity = bs->dsk.bitmap_granularity;
if (bs->dsk.meta_format >= BLOCKSTORE_META_FORMAT_V2)
{
hdr->data_csum_type = bs->dsk.data_csum_type;
hdr->csum_block_size = bs->dsk.csum_block_size;
}
if (bs->dsk.meta_format >= BLOCKSTORE_META_FORMAT_HEAP)
{
hdr->meta_area_size = bs->dsk.meta_area_size;
}
hdr->set_crc32c();
}
assert(bs->dsk.meta_format == BLOCKSTORE_META_FORMAT_HEAP);
blockstore_meta_header_v3_t *hdr = (blockstore_meta_header_v3_t *)bs->meta_superblock;
hdr->zero = 0;
hdr->magic = BLOCKSTORE_META_MAGIC_V1;
hdr->version = bs->dsk.meta_format;
hdr->meta_block_size = bs->dsk.meta_block_size;
hdr->data_block_size = bs->dsk.data_block_size;
hdr->bitmap_granularity = bs->dsk.bitmap_granularity;
hdr->completed_lsn = 0;
hdr->data_csum_type = bs->dsk.data_csum_type;
hdr->csum_block_size = bs->dsk.csum_block_size;
hdr->meta_area_size = bs->dsk.meta_area_size;
hdr->set_crc32c();
if (bs->readonly)
{
printf("Skipping metadata initialization because blockstore is readonly\n");
@@ -98,21 +92,8 @@ resume_1:
else
{
printf("Initializing metadata area\n");
GET_SQE();
last_read_offset = 0;
data->iov = (struct iovec){ bs->meta_superblock, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
io_uring_prep_writev(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset);
bs->ringloop->submit();
submitted++;
resume_2:
if (submitted > 0)
{
wait_state = 2;
return 1;
}
zero_on_init = true;
}
zero_on_init = true;
}
else
{
@@ -163,7 +144,7 @@ resume_1:
hdr->header_csum = csum;
}
bs->heap->start_load(((blockstore_meta_header_v3_t *)bs->meta_superblock)->completed_lsn);
if (bs->dsk.inmemory_journal)
if (bs->dsk.inmemory_journal && !zero_on_init)
{
// Read buffer area
printf("Reading buffered data\n");
@@ -175,7 +156,7 @@ resume_1:
bs->buffer_area + md_offset,
(size_t)(bs->dsk.journal_len - md_offset < bs->metadata_buf_size ? bs->dsk.journal_len - md_offset : bs->metadata_buf_size),
};
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "read buffer area"); };
io_uring_prep_readv(sqe, bs->dsk.journal_fd, &data->iov, 1, bs->dsk.journal_offset + md_offset);
md_offset += data->iov.iov_len;
submitted++;
@@ -194,7 +175,7 @@ resume_3:
next_offset = md_offset;
// Read the rest of the metadata
resume_4:
if (next_offset < bs->dsk.meta_area_size && submitted == 0)
if (next_offset < bs->dsk.meta_area_size && submitted == 0 && (!zero_on_init || !bs->readonly))
{
// Submit one read
for (int i = 0; i < 2; i++)
@@ -211,12 +192,15 @@ resume_4:
GET_SQE();
assert(bufs[i].size <= 0x7fffffff);
data->iov = { bufs[i].buf, (size_t)bufs[i].size };
data->callback = [this, i](ring_data_t *data) { handle_event(data, i); };
if (!zero_on_init)
{
data->callback = [this, i](ring_data_t *data) { handle_event(data, i, "read metadata"); };
io_uring_prep_readv(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset + bufs[i].offset);
}
else
{
// Fill metadata with empty block pattern
data->callback = [this, i](ring_data_t *data) { handle_event(data, i, "clear metadata"); };
memset(bufs[i].buf, 0, bufs[i].size);
for (uint64_t o = 0; o < bufs[i].size; o += bs->dsk.meta_block_size)
bs->heap->fill_block_empty_space(bufs[i].buf + o, 0);
@@ -232,11 +216,14 @@ resume_4:
if (bufs[i].state == INIT_META_READ_DONE)
{
// Handle result
uint64_t loaded = 0;
int r = bs->heap->load_blocks(bufs[i].offset-bs->dsk.meta_block_size, bufs[i].size, bufs[i].buf, bs->skip_corrupted_meta_entries, loaded);
if (r != 0)
exit(1);
entries_loaded += loaded;
if (!zero_on_init)
{
uint64_t loaded = 0;
int r = bs->heap->load_blocks(bufs[i].offset-bs->dsk.meta_block_size, bufs[i].size, bufs[i].buf, bs->skip_corrupted_meta_entries, loaded);
if (r != 0)
exit(1);
entries_loaded += loaded;
}
bufs[i].state = 0;
bs->ringloop->wakeup();
}
@@ -246,7 +233,7 @@ resume_4:
wait_state = 4;
return 1;
}
// metadata read finished
// metadata read/clear finished
bs->heap->finish_load();
printf("Metadata entries loaded: %ju, rechecking unfinished writes and garbage entries\n", entries_loaded);
// asynchronous recheck
@@ -329,13 +316,14 @@ resume_9:
}
free(metadata_buffer);
metadata_buffer = NULL;
do_fsync:
if (!bs->dsk.disable_meta_fsync && !bs->readonly)
{
GET_SQE();
io_uring_prep_fsync(sqe, bs->dsk.meta_fd, IORING_FSYNC_DATASYNC);
last_read_offset = 0;
data->iov = { 0 };
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "fsync metadata"); };
submitted++;
bs->ringloop->submit();
resume_5:
@@ -345,6 +333,27 @@ resume_9:
return 1;
}
}
if (zero_on_init && !header_written && !bs->readonly)
{
GET_SQE();
header_written = true;
last_read_offset = 0;
data->iov = (struct iovec){ bs->meta_superblock, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "write metadata header"); };
io_uring_prep_writev(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset);
bs->ringloop->submit();
submitted++;
resume_2:
if (submitted > 0)
{
wait_state = 2;
return 1;
}
if (!bs->dsk.disable_meta_fsync)
{
goto do_fsync;
}
}
printf("Loading finished. Data used: %ju / %ju bytes (%s / %s)\n",
bs->heap->get_data_used_space(), bs->dsk.block_count * bs->dsk.data_block_size,
format_size(bs->heap->get_data_used_space()).c_str(),
+2 -1
View File
@@ -17,6 +17,7 @@ class blockstore_init_meta
int wait_state = 0;
int wait_count = 0;
bool zero_on_init = false;
bool header_written = false;
void *metadata_buffer = NULL;
blockstore_init_meta_buf bufs[2] = {};
int submitted = 0;
@@ -29,7 +30,7 @@ class blockstore_init_meta
std::vector<uint32_t> recheck_mod;
int i = 0, j = 0;
bool handle_meta_block(uint8_t *buf, uint64_t count, uint64_t done_cnt);
void handle_event(ring_data_t *data, int buf_num);
void handle_event(ring_data_t *data, int buf_num, const char *op);
public:
blockstore_init_meta(blockstore_impl_t *bs);
int loop();
+113
View File
@@ -0,0 +1,113 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 (see README.md for details)
#include "blockstore_mock.h"
blockstore_mock_t::blockstore_mock_t(const blockstore_config_t & config)
{
}
void blockstore_mock_t::parse_config(blockstore_config_t & config)
{
}
void* blockstore_mock_t::reshard_start(pool_id_t pool, uint32_t pg_count, uint32_t pg_stripe_size, uint64_t chunk_limit)
{
return NULL;
}
bool blockstore_mock_t::reshard_continue(void *reshard_state, uint64_t chunk_limit)
{
return true;
}
void blockstore_mock_t::loop()
{
}
bool blockstore_mock_t::is_started()
{
return true;
}
bool blockstore_mock_t::is_stalled()
{
return false;
}
bool blockstore_mock_t::is_safe_to_stop()
{
return true;
}
void blockstore_mock_t::enqueue_op(blockstore_op_t *op)
{
}
int blockstore_mock_t::read_bitmap(object_id oid, uint64_t target_version, void *bitmap, uint64_t *result_version)
{
return -EIO;
}
const std::map<uint64_t, uint64_t> & blockstore_mock_t::get_inode_space_stats()
{
return inode_space;
}
void blockstore_mock_t::set_no_inode_stats(const std::vector<uint64_t> & pool_ids)
{
}
void blockstore_mock_t::dump_diagnostics()
{
}
std::string blockstore_mock_t::get_op_diag(blockstore_op_t *op)
{
return "";
}
uint32_t blockstore_mock_t::get_block_size()
{
return block_size;
}
uint64_t blockstore_mock_t::get_block_count()
{
return block_count;
}
uint64_t blockstore_mock_t::get_free_block_count()
{
return block_count;
}
uint64_t blockstore_mock_t::get_journal_size()
{
return 32*1024*1024;
}
uint32_t blockstore_mock_t::get_bitmap_granularity()
{
return bitmap_granularity;
}
uint64_t blockstore_mock_t::get_live_entries()
{
return 0;
}
uint64_t blockstore_mock_t::get_live_memory()
{
return 0;
}
uint64_t blockstore_mock_t::get_garbage_entries()
{
return 0;
}
uint64_t blockstore_mock_t::get_garbage_memory()
{
return 0;
}
+39
View File
@@ -0,0 +1,39 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 (see README.md for details)
#pragma once
#include "blockstore.h"
class blockstore_mock_t: public blockstore_i
{
public:
uint32_t block_size = 128*1024;
uint32_t bitmap_granularity = 4096;
uint64_t block_count = 100*1024*8;
std::map<uint64_t, uint64_t> inode_space;
blockstore_mock_t(const blockstore_config_t & config);
void parse_config(blockstore_config_t & config) override;
void* reshard_start(pool_id_t pool, uint32_t pg_count, uint32_t pg_stripe_size, uint64_t chunk_limit) override;
bool reshard_continue(void *reshard_state, uint64_t chunk_limit) override;
void loop() override;
bool is_started() override;
bool is_stalled() override;
bool is_safe_to_stop() override;
void enqueue_op(blockstore_op_t *op) override;
int read_bitmap(object_id oid, uint64_t target_version, void *bitmap, uint64_t *result_version = NULL) override;
const std::map<uint64_t, uint64_t> & get_inode_space_stats() override;
void set_no_inode_stats(const std::vector<uint64_t> & pool_ids) override;
void dump_diagnostics() override;
std::string get_op_diag(blockstore_op_t *op) override;
uint32_t get_block_size() override;
uint64_t get_block_count() override;
uint64_t get_free_block_count() override;
uint64_t get_journal_size() override;
uint32_t get_bitmap_granularity() override;
uint64_t get_live_entries() override;
uint64_t get_live_memory() override;
uint64_t get_garbage_entries() override;
uint64_t get_garbage_memory() override;
};
+1 -5
View File
@@ -28,6 +28,7 @@ int blockstore_impl_t::dequeue_stable(blockstore_op_t *op)
FINISH_OP(op);
return 2;
}
priv->modified_block2 = UINT32_MAX;
int res = op->opcode == BS_OP_STABLE
? heap->add_commit(obj, v[priv->stab_pos].version, &priv->modified_block2)
: heap->add_rollback(obj, v[priv->stab_pos].version, &priv->modified_block2);
@@ -52,11 +53,6 @@ int blockstore_impl_t::dequeue_stable(blockstore_op_t *op)
FINISH_OP(op);
return 2;
}
if (priv->modified_block2 != UINT32_MAX)
{
priv->stab_pos--;
goto resume_1;
}
priv->wait_for = WAIT_COMPACTION;
priv->wait_detail = heap->get_compacted_count();
flusher->request_trim();
+12 -5
View File
@@ -29,9 +29,12 @@ bool blockstore_impl_t::has_unsynced()
bool blockstore_impl_t::submit_fsyncs(int & wait_count)
{
int n = (unsynced_meta_write_count > 0 && !dsk.disable_meta_fsync) +
(unsynced_buffer_write_count > 0 && !dsk.disable_journal_fsync && dsk.journal_fd != dsk.meta_fd) +
(unsynced_data_write_count > 0 && !dsk.disable_data_fsync && dsk.data_fd != dsk.meta_fd && dsk.data_fd != dsk.journal_fd);
int n = (unsynced_meta_write_count > 0 && !dsk.disable_meta_fsync ? 1 : 0) +
(unsynced_buffer_write_count > 0 && !dsk.disable_journal_fsync &&
(!unsynced_meta_write_count || dsk.journal_fd != dsk.meta_fd) ? 1 : 0) +
(unsynced_data_write_count > 0 && !dsk.disable_data_fsync &&
(!unsynced_meta_write_count || dsk.data_fd != dsk.meta_fd) &&
(!unsynced_buffer_write_count || dsk.data_fd != dsk.journal_fd) ? 1 : 0);
if (ringloop->space_left() < n)
{
return false;
@@ -60,7 +63,8 @@ bool blockstore_impl_t::submit_fsyncs(int & wait_count)
data->callback = cb;
wait_count++;
}
if (unsynced_buffer_write_count > 0 && !dsk.disable_journal_fsync && dsk.meta_fd != dsk.journal_fd)
if (unsynced_buffer_write_count > 0 && !dsk.disable_journal_fsync &&
(!unsynced_meta_write_count || dsk.journal_fd != dsk.meta_fd))
{
// fsync buffer
io_uring_sqe *sqe = get_sqe();
@@ -71,7 +75,9 @@ bool blockstore_impl_t::submit_fsyncs(int & wait_count)
data->callback = cb;
wait_count++;
}
if (unsynced_data_write_count > 0 && !dsk.disable_data_fsync && dsk.data_fd != dsk.meta_fd && dsk.data_fd != dsk.journal_fd)
if (unsynced_data_write_count > 0 && !dsk.disable_data_fsync &&
(!unsynced_meta_write_count || dsk.data_fd != dsk.meta_fd) &&
(!unsynced_buffer_write_count || dsk.data_fd != dsk.journal_fd))
{
// fsync data
io_uring_sqe *sqe = get_sqe();
@@ -109,6 +115,7 @@ int blockstore_impl_t::do_sync(blockstore_op_t *op, int base_state)
PRIV(op)->lsn = heap->get_completed_lsn();
if (!submit_fsyncs(PRIV(op)->pending_ops))
{
PRIV(op)->lsn = 0;
PRIV(op)->wait_detail = 1;
PRIV(op)->wait_for = WAIT_SQE;
return 0;
+35 -28
View File
@@ -180,13 +180,16 @@ enospc:
data->callback = [this, op](ring_data_t *data) { handle_write_event(data, op); };
assert(loc+op->offset+op->len <= dsk.block_count*dsk.data_block_size);
io_uring_prep_writev(sqe, dsk.data_fd, &data->iov, 1, dsk.data_offset + loc + op->offset);
if (!dsk.disable_data_fsync)
{
// use PRIV->lsn for fsync_data_id
PRIV(op)->lsn = ++data_fsync_next;
data_fsyncs.push_back(false);
}
PRIV(op)->pending_ops++;
write_iodepth++;
if (PRIV(op)->write_type == BS_HEAP_BIG_WRITE)
{
PRIV(op)->op_state = 1;
inflight_big++;
}
else
PRIV(op)->op_state = 3;
}
@@ -297,8 +300,6 @@ again:
goto resume_10;
else if (op_state == 11)
goto resume_11;
else if (op_state == 12)
goto resume_12;
else
{
// In progress
@@ -317,38 +318,44 @@ again:
resume_2:
// We must fsync all big writes to avoid complex write workflows
// It's OK for all HDDs and for server SSDs, but slightly worse for desktop SSDs
inflight_big--;
if (!dsk.disable_data_fsync)
{
// fsync data in a batch
resume_11:
if (inflight_big > 0)
// Mark our data write as completed and advance data_fsync_cur
data_fsyncs[PRIV(op)->lsn - data_fsync_cur - 1] = true;
while (data_fsyncs.size() > 0 && data_fsyncs.front())
{
data_fsyncs.pop_front();
data_fsync_cur++;
}
PRIV(op)->op_state = 11;
// Then wait for all other data writes currently in progress to do less fsync calls
// I.e. to fsync data in batches
PRIV(op)->lsn = data_fsync_cur + data_fsyncs.size();
resume_11:
if (data_fsync_cur < PRIV(op)->lsn)
{
PRIV(op)->op_state = 11;
return 1;
}
if (fsyncing_data)
if (PRIV(op)->lsn > data_fsync_sent)
{
resume_12:
if (fsyncing_data)
BS_SUBMIT_GET_SQE(sqe, data);
io_uring_prep_fsync(sqe, dsk.data_fd, IORING_FSYNC_DATASYNC);
data->iov = { 0 };
data->callback = [this, op, fs = data_fsync_cur](ring_data_t *data)
{
PRIV(op)->op_state = 12;
return 1;
}
goto resume_4;
if (fs > data_fsync_done)
{
data_fsync_done = fs;
ringloop->wakeup();
}
};
data_fsync_sent = data_fsync_cur;
}
fsyncing_data = true;
BS_SUBMIT_GET_SQE(sqe, data);
io_uring_prep_fsync(sqe, dsk.data_fd, IORING_FSYNC_DATASYNC);
data->iov = { 0 };
data->callback = [this, op](ring_data_t *data)
if (PRIV(op)->lsn > data_fsync_done)
{
fsyncing_data = false;
handle_write_event(data, op);
};
PRIV(op)->pending_ops++;
PRIV(op)->op_state = 3;
return 1;
return 1;
}
PRIV(op)->lsn = 0;
}
resume_4:
{
+7 -4
View File
@@ -171,7 +171,7 @@ void multilist_alloc_t::print()
printf("\n");
}
void multilist_alloc_t::use(uint32_t pos, uint32_t size)
bool multilist_alloc_t::use(uint32_t pos, uint32_t size)
{
assert(pos+size <= count && size > 0);
if (sizes[pos] <= 0)
@@ -182,7 +182,8 @@ void multilist_alloc_t::use(uint32_t pos, uint32_t size)
else
while (start > 0 && !sizes[start])
start--;
assert(sizes[start] >= size);
if (sizes[start] < size+(pos-start))
return false;
use_full(start);
uint32_t full = sizes[start];
sizes[pos-1] = -pos+start;
@@ -199,7 +200,8 @@ void multilist_alloc_t::use(uint32_t pos, uint32_t size)
}
else
{
assert(sizes[pos] >= size);
if (sizes[pos] < size)
return false;
use_full(pos);
if (sizes[pos] > size)
{
@@ -214,12 +216,13 @@ void multilist_alloc_t::use(uint32_t pos, uint32_t size)
#ifdef MULTILIST_TRACE
print();
#endif
return true;
}
void multilist_alloc_t::use_full(uint32_t pos)
{
uint32_t prevsize = sizes[pos];
assert(prevsize);
assert(prevsize > 0);
assert(nexts[pos]);
uint32_t pi = (prevsize < maxn ? prevsize : maxn)-1;
if (heads[pi] == pos+1)
+1 -1
View File
@@ -17,7 +17,7 @@ struct multilist_alloc_t
bool is_free(uint32_t pos);
uint32_t find(uint32_t size);
void use_full(uint32_t pos);
void use(uint32_t pos, uint32_t size);
bool use(uint32_t pos, uint32_t size);
void do_free(uint32_t pos);
void free(uint32_t pos);
void verify();
+46 -21
View File
@@ -71,6 +71,11 @@ bool journal_flusher_t::is_active()
return active_flushers > 0 || dequeuing;
}
size_t journal_flusher_t::get_queue_size()
{
return flush_queue.size();
}
void journal_flusher_t::loop()
{
target_flusher_count = bs->write_iodepth*2;
@@ -384,6 +389,7 @@ stop_flusher:
wait_state = 0;
return true;
}
copy_count = 0;
try_trim = true;
cur.oid = flusher->flush_queue.front();
cur.version = flusher->flush_versions[cur.oid];
@@ -511,6 +517,31 @@ resume_2:
{
uo_it->second.was_changed = true;
}
if (!bs->journal.inmemory)
{
// Verify journaled data checksums (but not COALESCED)
for (it = v.begin(); it != v.end(); it++)
{
if (it->copy_flags == COPY_BUF_JOURNAL)
{
iovec iov = { .iov_base = it->buf, .iov_len = it->len };
bs->verify_journal_checksums(
it->csum_buf, it->offset, &iov, 1,
[&](uint32_t bad_block, uint32_t calc_csum, uint32_t stored_csum)
{
printf(
"Checksum mismatch in object %jx:%jx v%ju in journal at 0x%jx, checksum block #%u: got %08x, expected %08x\n",
cur.oid.inode, cur.oid.stripe, cur.version, it->disk_offset,
bad_block / bs->dsk.csum_block_size, calc_csum, stored_csum
);
bad_block += it->offset;
assert(!(bad_block % bs->dsk.csum_block_size) && bad_block < bs->dsk.data_block_size);
mangle_csum_blocks.insert(bad_block);
}
);
}
}
}
}
// Submit data writes
for (it = v.begin(); it != v.end(); it++)
@@ -634,6 +665,7 @@ resume_2:
}
// All done
flusher->active_flushers--;
copy_count = 0; // used by is_mutated()...
wait_state = 0;
goto resume_0;
}
@@ -815,35 +847,21 @@ bool journal_flusher_co::clear_incomplete_csum_block_bits(int wait_base)
bs->verify_padded_checksums(new_clean_bitmap, new_clean_bitmap + 2*bs->dsk.clean_entry_bitmap_size,
v[i].offset, &iov, 1, [&](uint32_t bad_block, uint32_t calc_csum, uint32_t stored_csum)
{
printf("Checksum mismatch in object %jx:%jx v%ju in data area at offset 0x%jx+0x%x: got %08x, expected %08x\n",
printf("Checksum mismatch in object %jx:%jx v%ju in data area at offset 0x%jx+0x%x during flush: got %08x, expected %08x\n",
cur.oid.inode, cur.oid.stripe, old_clean_ver, old_clean_loc, bad_block, calc_csum, stored_csum);
for (uint32_t j = 0; j < bs->dsk.csum_block_size; j += bs->dsk.bitmap_granularity)
{
// Simplest method of mangling: flip one byte in every sector
((uint8_t*)v[i].buf)[j+bad_block-v[i].offset] ^= 0xff;
}
assert(!(bad_block % bs->dsk.csum_block_size) && bad_block < bs->dsk.data_block_size);
mangle_csum_blocks.insert(bad_block);
});
}
else
{
bs->verify_journal_checksums(v[i].csum_buf, v[i].offset, &iov, 1, [&](uint32_t bad_block, uint32_t calc_csum, uint32_t stored_csum)
{
printf("Checksum mismatch in object %jx:%jx v%ju in journal at offset 0x%jx+0x%x (block offset 0x%jx): got %08x, expected %08x\n",
printf("Checksum mismatch in object %jx:%jx v%ju in journal at offset 0x%jx+0x%x (block offset 0x%jx) during flush: got %08x, expected %08x\n",
cur.oid.inode, cur.oid.stripe, old_clean_ver,
v[i].disk_offset, bad_block, v[i].offset, calc_csum, stored_csum);
bad_block += (v[i].offset/bs->dsk.csum_block_size) * bs->dsk.csum_block_size;
uint32_t bad_block_end = bad_block + bs->dsk.csum_block_size + (v[i].offset/bs->dsk.csum_block_size) * bs->dsk.csum_block_size;
if (bad_block < v[i].offset)
bad_block = v[i].offset;
if (bad_block_end > v[i].offset+v[i].len)
bad_block_end = v[i].offset+v[i].len;
bad_block -= v[i].offset;
bad_block_end -= v[i].offset;
for (uint32_t j = bad_block; j < bad_block_end; j += bs->dsk.bitmap_granularity)
{
// Simplest method of mangling: flip one byte in every sector
((uint8_t*)v[i].buf)[j] ^= 0xff;
}
assert(!(bad_block % bs->dsk.csum_block_size) && bad_block < bs->dsk.data_block_size);
mangle_csum_blocks.insert(bad_block);
});
}
}
@@ -952,6 +970,11 @@ void journal_flusher_co::calc_block_checksums(uint32_t *new_data_csums, bool ski
}
// `v` should contain aligned items, possibly split into pieces
assert(!block_done);
for (uint32_t mangle_block: mangle_csum_blocks)
{
// Flip 1 bit
new_data_csums[mangle_block / bs->dsk.csum_block_size] ^= 1;
}
}
void journal_flusher_co::scan_dirty()
@@ -1088,7 +1111,8 @@ void journal_flusher_co::scan_dirty()
last--;
read_to_fill_incomplete = bs->fill_partial_checksum_blocks(
v, fulfilled, bmp_ptr, NULL, false, NULL, v[0].offset/bs->dsk.csum_block_size * bs->dsk.csum_block_size,
((v[last].offset+v[last].len-1) / bs->dsk.csum_block_size + 1) * bs->dsk.csum_block_size
((v[last].offset+v[last].len-1) / bs->dsk.csum_block_size + 1) * bs->dsk.csum_block_size,
0, bs->dsk.data_block_size
);
}
else if (fill_incomplete && clean_init_bitmap)
@@ -1118,6 +1142,7 @@ bool journal_flusher_co::read_dirty(int wait_base)
if (wait_state == wait_base) goto resume_0;
else if (wait_state == wait_base+1) goto resume_1;
wait_count = wait_journal_count = 0;
mangle_csum_blocks.clear();
if (bs->journal.inmemory && !read_to_fill_incomplete)
{
// Happy path: nothing to read :)
+2
View File
@@ -66,6 +66,7 @@ class journal_flusher_co
uint64_t clean_bitmap_offset, clean_bitmap_len;
uint8_t *clean_init_dyn_ptr;
uint8_t *new_clean_bitmap;
std::unordered_set<uint32_t> mangle_csum_blocks;
uint64_t new_trim_pos;
@@ -123,6 +124,7 @@ public:
void loop();
bool is_trim_wanted() { return trim_wanted; }
bool is_active();
size_t get_queue_size();
void mark_trim_possible();
void request_trim();
void release_trim();
+7 -1
View File
@@ -6,11 +6,12 @@
namespace v1 {
blockstore_impl_t::blockstore_impl_t(blockstore_config_t & config, ring_loop_i *ringloop, timerfd_manager_t *tfd)
blockstore_impl_t::blockstore_impl_t(blockstore_config_t & config, ring_loop_i *ringloop, timerfd_manager_t *tfd, bool mock_mode)
{
assert(sizeof(blockstore_op_private_t) <= BS_OP_PRIVATE_DATA_SIZE);
this->tfd = tfd;
this->ringloop = ringloop;
dsk.mock_mode = mock_mode;
ring_consumer.loop = [this]() { loop(); };
ringloop->register_consumer(&ring_consumer);
initialized = 0;
@@ -35,6 +36,11 @@ blockstore_impl_t::blockstore_impl_t(blockstore_config_t & config, ring_loop_i *
blockstore_impl_t::~blockstore_impl_t()
{
for (auto& obj: dirty_db)
{
if (obj.second.dyn_data)
free(obj.second.dyn_data);
}
delete data_alloc;
delete flusher;
if (zero_object)
+8 -3
View File
@@ -30,6 +30,8 @@
//#define BLOCKSTORE_DEBUG
struct bs_test_t;
namespace v1 {
#include "journal.h"
@@ -96,7 +98,7 @@ struct blockstore_op_private_t
int op_state;
// Read
uint64_t clean_block_used;
uint64_t clean_loc_used;
std::vector<copy_buffer_t> read_vec;
// Sync, write
@@ -122,6 +124,7 @@ typedef uint64_t pool_pg_id_t;
class blockstore_impl_t: public blockstore_i
{
friend struct ::bs_test_t;
blockstore_disk_t dsk;
/******* OPTIONS *******/
@@ -220,6 +223,7 @@ class blockstore_impl_t: public blockstore_i
// Read
int dequeue_read(blockstore_op_t *read_op);
void release_clean(blockstore_op_t *op);
void find_holes(std::vector<copy_buffer_t> & read_vec, uint32_t item_start, uint32_t item_end,
std::function<int(int, bool, uint32_t, uint32_t)> callback);
int fulfill_read(blockstore_op_t *read_op,
@@ -230,7 +234,8 @@ class blockstore_impl_t: public blockstore_i
uint8_t *clean_entry_bitmap, int *dyn_data,
uint32_t item_start, uint32_t item_end, uint64_t clean_loc, uint64_t clean_ver);
int fill_partial_checksum_blocks(std::vector<copy_buffer_t> & rv, uint64_t & fulfilled,
uint8_t *clean_entry_bitmap, int *dyn_data, bool from_journal, uint8_t *read_buf, uint64_t read_offset, uint64_t read_end);
uint8_t *clean_entry_bitmap, int *dyn_data, bool from_journal, uint8_t *read_buf,
uint32_t read_offset, uint32_t read_end, uint32_t item_start, uint32_t item_end);
int pad_journal_read(std::vector<copy_buffer_t> & rv, copy_buffer_t & cp,
uint64_t dirty_offset, uint64_t dirty_end, uint64_t dirty_loc, uint8_t *csum_ptr, int *dyn_data,
uint64_t offset, uint64_t submit_len, uint64_t & blk_begin, uint64_t & blk_end, uint8_t* & blk_buf);
@@ -281,7 +286,7 @@ class blockstore_impl_t: public blockstore_i
public:
blockstore_impl_t(blockstore_config_t & config, ring_loop_i *ringloop, timerfd_manager_t *tfd);
blockstore_impl_t(blockstore_config_t & config, ring_loop_i *ringloop, timerfd_manager_t *tfd, bool mock_mode = false);
~blockstore_impl_t();
void parse_config(blockstore_config_t & config);
+83 -68
View File
@@ -1,6 +1,7 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 (see README.md for details)
#include "str_util.h"
#include "impl.h"
#include "internal.h"
@@ -30,14 +31,15 @@ blockstore_init_meta::blockstore_init_meta(blockstore_impl_t *bs)
this->bs = bs;
}
void blockstore_init_meta::handle_event(ring_data_t *data, int buf_num)
void blockstore_init_meta::handle_event(ring_data_t *data, int buf_num, const char *op)
{
if (data->res < 0)
if (data->res != data->iov.iov_len)
{
throw std::runtime_error(
std::string("read metadata failed at offset ") + std::to_string(buf_num >= 0 ? bufs[buf_num].offset : last_read_offset) +
std::string(": ") + strerror(-data->res)
);
throw std::runtime_error(strprintf(
"%s failed at offset %ju: got %s (code %d), but expected %zu",
op, (buf_num >= 0 ? bufs[buf_num].offset : last_read_offset), strerror(-data->res),
data->res, data->iov.iov_len
));
}
if (buf_num >= 0)
{
@@ -65,10 +67,11 @@ int blockstore_init_meta::loop()
if (!metadata_buffer)
throw std::runtime_error("Failed to allocate metadata read buffer");
// Read superblock
hdr = (blockstore_meta_header_v2_t *)memalign_or_die(MEM_ALIGNMENT, bs->dsk.meta_block_size);
GET_SQE();
last_read_offset = 0;
data->iov = { metadata_buffer, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
data->iov = { hdr, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "read metadata header"); };
io_uring_prep_readv(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset);
bs->ringloop->submit();
submitted++;
@@ -78,24 +81,8 @@ resume_1:
wait_state = 1;
return 1;
}
if (iszero((uint64_t*)metadata_buffer, bs->dsk.meta_block_size / sizeof(uint64_t)))
if (iszero((uint64_t*)hdr, bs->dsk.meta_block_size / sizeof(uint64_t)))
{
{
blockstore_meta_header_v2_t *hdr = (blockstore_meta_header_v2_t *)metadata_buffer;
hdr->zero = 0;
hdr->magic = BLOCKSTORE_META_MAGIC_V1;
hdr->version = bs->dsk.meta_format;
hdr->meta_block_size = bs->dsk.meta_block_size;
hdr->data_block_size = bs->dsk.data_block_size;
hdr->bitmap_granularity = bs->dsk.bitmap_granularity;
if (bs->dsk.meta_format >= BLOCKSTORE_META_FORMAT_V2)
{
hdr->data_csum_type = bs->dsk.data_csum_type;
hdr->csum_block_size = bs->dsk.csum_block_size;
hdr->header_csum = 0;
hdr->header_csum = crc32c(0, hdr, sizeof(*hdr));
}
}
if (bs->readonly)
{
printf("Skipping metadata initialization because blockstore is readonly\n");
@@ -103,25 +90,11 @@ resume_1:
else
{
printf("Initializing metadata area\n");
GET_SQE();
last_read_offset = 0;
data->iov = (struct iovec){ metadata_buffer, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
io_uring_prep_writev(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset);
bs->ringloop->submit();
submitted++;
resume_3:
if (submitted > 0)
{
wait_state = 3;
return 1;
}
zero_on_init = true;
}
zero_on_init = true;
}
else
{
blockstore_meta_header_v2_t *hdr = (blockstore_meta_header_v2_t *)metadata_buffer;
if (hdr->zero != 0 || hdr->magic != BLOCKSTORE_META_MAGIC_V1 || hdr->version < BLOCKSTORE_META_FORMAT_V1)
{
printf(
@@ -223,12 +196,15 @@ resume_2:
GET_SQE();
assert(bufs[i].size <= 0x7fffffff);
data->iov = { bufs[i].buf, (size_t)bufs[i].size };
data->callback = [this, i](ring_data_t *data) { handle_event(data, i); };
if (!zero_on_init)
{
data->callback = [this, i](ring_data_t *data) { handle_event(data, i, "read metadata"); };
io_uring_prep_readv(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset + bufs[i].offset);
}
else
{
// Fill metadata with zeroes
data->callback = [this, i](ring_data_t *data) { handle_event(data, i, "clear metadata"); };
memset(data->iov.iov_base, 0, data->iov.iov_len);
io_uring_prep_writev(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset + bufs[i].offset);
}
@@ -256,7 +232,7 @@ resume_2:
GET_SQE();
assert(bufs[i].size <= 0x7fffffff);
data->iov = { bufs[i].buf, (size_t)bufs[i].size };
data->callback = [this, i](ring_data_t *data) { handle_event(data, i); };
data->callback = [this, i](ring_data_t *data) { handle_event(data, i, "write metadata"); };
io_uring_prep_writev(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset + bufs[i].offset);
bs->ringloop->submit();
bufs[i].state = INIT_META_WRITING;
@@ -285,7 +261,7 @@ resume_2:
GET_SQE();
last_read_offset = (1+next_offset)*bs->dsk.meta_block_size;
data->iov = { metadata_buffer, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "read metadata"); };
io_uring_prep_readv(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset + (1+next_offset)*bs->dsk.meta_block_size);
bs->ringloop->submit();
submitted++;
@@ -302,7 +278,7 @@ resume_5:
}
GET_SQE();
data->iov = { metadata_buffer, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "write metadata"); };
io_uring_prep_writev(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset + (1+next_offset)*bs->dsk.meta_block_size);
bs->ringloop->submit();
submitted++;
@@ -317,27 +293,64 @@ resume_6:
}
// metadata read finished
printf("Metadata entries loaded: %ju, free blocks: %ju / %ju\n", entries_loaded, bs->data_alloc->get_free_count(), bs->dsk.block_count);
if (zero_on_init && !bs->readonly)
{
do_fsync:
if (!bs->disable_meta_fsync)
{
GET_SQE();
io_uring_prep_fsync(sqe, bs->dsk.meta_fd, IORING_FSYNC_DATASYNC);
last_read_offset = 0;
data->iov = { 0 };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "fsync metadata"); };
submitted++;
bs->ringloop->submit();
resume_4:
if (submitted > 0)
{
wait_state = 4;
return 1;
}
}
if (!header_written)
{
GET_SQE();
hdr->zero = 0;
hdr->magic = BLOCKSTORE_META_MAGIC_V1;
hdr->version = bs->dsk.meta_format;
hdr->meta_block_size = bs->dsk.meta_block_size;
hdr->data_block_size = bs->dsk.data_block_size;
hdr->bitmap_granularity = bs->dsk.bitmap_granularity;
if (bs->dsk.meta_format >= BLOCKSTORE_META_FORMAT_V2)
{
hdr->data_csum_type = bs->dsk.data_csum_type;
hdr->csum_block_size = bs->dsk.csum_block_size;
hdr->header_csum = 0;
hdr->header_csum = crc32c(0, hdr, sizeof(*hdr));
}
header_written = true;
last_read_offset = 0;
data->iov = (struct iovec){ hdr, (size_t)bs->dsk.meta_block_size };
data->callback = [this](ring_data_t *data) { handle_event(data, -1, "write metadata header"); };
io_uring_prep_writev(sqe, bs->dsk.meta_fd, &data->iov, 1, bs->dsk.meta_offset);
bs->ringloop->submit();
submitted++;
resume_3:
if (submitted > 0)
{
wait_state = 3;
return 1;
}
goto do_fsync;
}
}
if (!bs->inmemory_meta)
{
free(metadata_buffer);
metadata_buffer = NULL;
}
if (zero_on_init && !bs->disable_meta_fsync)
{
GET_SQE();
io_uring_prep_fsync(sqe, bs->dsk.meta_fd, IORING_FSYNC_DATASYNC);
last_read_offset = 0;
data->iov = { 0 };
data->callback = [this](ring_data_t *data) { handle_event(data, -1); };
submitted++;
bs->ringloop->submit();
resume_4:
if (submitted > 0)
{
wait_state = 4;
return 1;
}
}
free(hdr);
hdr = NULL;
return 0;
}
@@ -345,6 +358,8 @@ bool blockstore_init_meta::handle_meta_block(uint8_t *buf, uint64_t entries_per_
{
bool updated = false;
uint64_t max_i = entries_per_block;
if (done_cnt > bs->dsk.block_count)
return false;
if (max_i > bs->dsk.block_count-done_cnt)
max_i = bs->dsk.block_count-done_cnt;
for (uint64_t i = 0; i < max_i; i++)
@@ -455,21 +470,21 @@ blockstore_init_journal::blockstore_init_journal(blockstore_impl_t *bs)
};
}
void blockstore_init_journal::handle_event(ring_data_t *data1)
void blockstore_init_journal::handle_event(ring_data_t *data)
{
if (data1->res <= 0)
if (data->res != data->iov.iov_len)
{
throw std::runtime_error(
std::string("read journal failed at offset ") + std::to_string(journal_pos) +
std::string(": ") + strerror(-data1->res)
);
throw std::runtime_error(strprintf(
"read journal failed at offset %ju: got %s (code %d), but expected %zu",
journal_pos, strerror(-data->res), data->res, data->iov.iov_len
));
}
done.push_back({
.buf = submitted_buf,
.pos = journal_pos,
.len = (uint64_t)data1->res,
.len = (uint64_t)data->res,
});
journal_pos += data1->res;
journal_pos += data->res;
if (journal_pos >= bs->journal.len)
{
// Continue from the beginning
+3 -1
View File
@@ -16,7 +16,9 @@ class blockstore_init_meta
blockstore_impl_t *bs;
int wait_state = 0;
bool zero_on_init = false;
bool header_written = false;
void *metadata_buffer = NULL;
blockstore_meta_header_v2_t *hdr = NULL;
blockstore_init_meta_buf bufs[2] = {};
int submitted = 0;
struct io_uring_sqe *sqe;
@@ -29,7 +31,7 @@ class blockstore_init_meta
int i = 0, j = 0;
std::vector<uint64_t> entries_to_zero;
bool handle_meta_block(uint8_t *buf, uint64_t count, uint64_t done_cnt);
void handle_event(ring_data_t *data, int buf_num);
void handle_event(ring_data_t *data, int buf_num, const char *op);
public:
blockstore_init_meta(blockstore_impl_t *bs);
int loop();
+145 -54
View File
@@ -101,8 +101,8 @@ int blockstore_impl_t::fulfill_read(blockstore_op_t *read_op,
.copy_flags = COPY_BUF_JOURNAL|COPY_BUF_CSUM_FILL,
.offset = blk_begin,
.len = blk_end-blk_begin,
.csum_buf = (csum + (blk_begin/dsk.csum_block_size -
item_start/dsk.csum_block_size) * (dsk.data_csum_type & 0xFF)),
.csum_buf = (!csum ? NULL : (csum + (blk_begin/dsk.csum_block_size -
item_start/dsk.csum_block_size) * (dsk.data_csum_type & 0xFF))),
.dyn_data = dyn_data,
});
if (dyn_data)
@@ -134,7 +134,7 @@ int blockstore_impl_t::fulfill_read(blockstore_op_t *read_op,
// If we don't track it then we may IN THEORY read another object's data:
// submit read -> remove the object -> flush remove -> overwrite with another object -> finish read
// Very improbable, but possible
PRIV(read_op)->clean_block_used = 1;
PRIV(read_op)->clean_loc_used = UINT64_MAX;
}
rv.insert(rv.begin() + pos, el);
fulfilled += el.len;
@@ -167,7 +167,8 @@ uint8_t* blockstore_impl_t::get_clean_entry_bitmap(uint64_t block_loc, int offse
}
int blockstore_impl_t::fill_partial_checksum_blocks(std::vector<copy_buffer_t> & rv, uint64_t & fulfilled,
uint8_t *clean_entry_bitmap, int *dyn_data, bool from_journal, uint8_t *read_buf, uint64_t read_offset, uint64_t read_end)
uint8_t *clean_entry_bitmap, int *dyn_data, bool from_journal, uint8_t *read_buf,
uint32_t read_offset, uint32_t read_end, uint32_t item_start, uint32_t item_end)
{
if (read_end == read_offset)
return 0;
@@ -175,10 +176,38 @@ int blockstore_impl_t::fill_partial_checksum_blocks(std::vector<copy_buffer_t> &
read_buf -= read_offset;
uint32_t last_block = (read_end-1)/dsk.csum_block_size;
uint32_t start_block = read_offset/dsk.csum_block_size;
uint32_t item_start_block = item_start/dsk.csum_block_size;
uint32_t end_block = 0;
auto zero_range = [&](int pos, bool alloc, uint32_t cur_start, uint32_t cur_end)
{
if (alloc)
return 0;
copy_buffer_t el = {
.copy_flags = COPY_BUF_ZERO,
.offset = cur_start,
.len = cur_end-cur_start,
};
rv.insert(rv.begin() + pos, el);
if (read_buf)
memset(read_buf + el.offset - read_offset, 0, el.len);
fulfilled += el.len;
return 1;
};
if (read_offset < item_start)
{
// Zero-fill the beginning
find_holes(rv, read_offset, item_start, zero_range);
read_offset = item_start;
}
if (read_end > item_end)
{
// Zero-fill the end
find_holes(rv, item_end, read_end, zero_range);
read_end = item_end;
}
while (start_block <= last_block)
{
if (read_range_fulfilled(rv, fulfilled, read_buf, clean_entry_bitmap,
if (read_range_fulfilled(rv, fulfilled, read_buf, from_journal ? NULL : clean_entry_bitmap,
start_block*dsk.csum_block_size < read_offset ? read_offset : start_block*dsk.csum_block_size,
(start_block+1)*dsk.csum_block_size > read_end ? read_end : (start_block+1)*dsk.csum_block_size))
{
@@ -190,7 +219,7 @@ int blockstore_impl_t::fill_partial_checksum_blocks(std::vector<copy_buffer_t> &
// Find a sequence of checksum blocks required to be read
end_block = start_block;
while ((end_block+1)*dsk.csum_block_size < read_end &&
!read_range_fulfilled(rv, fulfilled, read_buf, clean_entry_bitmap,
!read_range_fulfilled(rv, fulfilled, read_buf, from_journal ? NULL : clean_entry_bitmap,
(end_block+1)*dsk.csum_block_size < read_offset ? read_offset : (end_block+1)*dsk.csum_block_size,
(end_block+2)*dsk.csum_block_size > read_end ? read_end : (end_block+2)*dsk.csum_block_size))
{
@@ -202,8 +231,10 @@ int blockstore_impl_t::fill_partial_checksum_blocks(std::vector<copy_buffer_t> &
.copy_flags = COPY_BUF_CSUM_FILL | (from_journal ? COPY_BUF_JOURNALED_BIG : 0),
.offset = start_block*dsk.csum_block_size,
.len = (end_block-start_block)*dsk.csum_block_size,
// save clean_entry_bitmap if we're reading clean data from the journal
.csum_buf = from_journal ? clean_entry_bitmap : NULL,
// save checksum reference if we're reading clean data from the journal
.csum_buf = from_journal
? clean_entry_bitmap + dsk.clean_entry_bitmap_size + (start_block-item_start_block)*(dsk.data_csum_type & 0xFF)
: NULL,
.dyn_data = dyn_data,
});
if (dyn_data)
@@ -226,6 +257,11 @@ bool blockstore_impl_t::read_range_fulfilled(std::vector<copy_buffer_t> & rv, ui
{
if (alloc)
return 0;
if (!clean_entry_bitmap)
{
all_done = false;
return 0;
}
int diff = 0;
uint32_t bmp_start = cur_start/dsk.bitmap_granularity;
uint32_t bmp_end = cur_end/dsk.bitmap_granularity;
@@ -323,7 +359,7 @@ bool blockstore_impl_t::read_checksum_block(blockstore_op_t *op, int rv_pos, uin
{
iov[n_iov++] = (struct iovec){ (uint8_t*)op->buf+cur_start-op->offset, lim_end-cur_start };
rv.insert(rv.begin() + pos, (copy_buffer_t){
.copy_flags = COPY_BUF_DATA,
.copy_flags = COPY_BUF_DATA|COPY_BUF_COALESCED,
.offset = cur_start,
.len = lim_end-cur_start,
});
@@ -361,10 +397,10 @@ bool blockstore_impl_t::read_checksum_block(blockstore_op_t *op, int rv_pos, uin
PRIV(op)->pending_ops++;
io_uring_prep_readv(sqe, submit_fd, iov + n_pos, n_cur, submit_offset + clean_loc + item_start + d_pos);
data->callback = [this, op](ring_data_t *data) { handle_read_event(data, op); };
if (n_pos > 0 || n_pos + IOV_MAX < n_iov)
if (n_pos > 0 || n_iov > IOV_MAX)
{
uint32_t d_len = 0;
for (int i = 0; i < IOV_MAX; i++)
for (int i = 0; i < n_cur; i++)
d_len += iov[n_pos+i].iov_len;
data->iov.iov_len = d_len;
d_pos += d_len;
@@ -376,7 +412,7 @@ bool blockstore_impl_t::read_checksum_block(blockstore_op_t *op, int rv_pos, uin
{
// Reads running parallel to flushes of the same clean block may read
// a mixture of old and new data. So we don't verify checksums for such blocks.
PRIV(op)->clean_block_used = 1;
PRIV(op)->clean_loc_used = UINT64_MAX;
}
return true;
}
@@ -402,7 +438,7 @@ int blockstore_impl_t::dequeue_read(blockstore_op_t *read_op)
}
uint64_t fulfilled = 0;
PRIV(read_op)->pending_ops = 0;
PRIV(read_op)->clean_block_used = 0;
PRIV(read_op)->clean_loc_used = 0;
auto & rv = PRIV(read_op)->read_vec;
uint64_t result_version = 0;
if (dirty_found)
@@ -515,26 +551,50 @@ int blockstore_impl_t::dequeue_read(blockstore_op_t *read_op)
return 2;
undo_read:
// need to wait. undo added requests, don't dequeue op
if (dsk.csum_block_size > dsk.bitmap_granularity)
release_clean(read_op);
for (auto & vec: rv)
{
for (auto & vec: rv)
if ((vec.copy_flags & COPY_BUF_CSUM_FILL) && vec.buf)
{
if ((vec.copy_flags & COPY_BUF_CSUM_FILL) && vec.buf)
{
free(vec.buf);
vec.buf = NULL;
}
if (vec.dyn_data && --(*vec.dyn_data) == 0) // refcount
{
free(vec.dyn_data);
vec.dyn_data = NULL;
}
free(vec.buf);
vec.buf = NULL;
}
if (vec.dyn_data && --(*vec.dyn_data) == 0) // refcount
{
free(vec.dyn_data);
vec.dyn_data = NULL;
}
}
rv.clear();
return 0;
}
void blockstore_impl_t::release_clean(blockstore_op_t *op)
{
if (PRIV(op)->clean_loc_used == UINT64_MAX)
{
PRIV(op)->clean_loc_used = 0;
}
if (PRIV(op)->clean_loc_used)
{
// Release clean data block
auto uo_it = used_clean_objects.find(PRIV(op)->clean_loc_used - 1);
if (uo_it != used_clean_objects.end())
{
uo_it->second.refs--;
if (uo_it->second.refs <= 0)
{
if (uo_it->second.was_freed)
{
data_alloc->set((PRIV(op)->clean_loc_used - 1) / dsk.data_block_size, false);
}
used_clean_objects.erase(uo_it);
}
}
PRIV(op)->clean_loc_used = 0;
}
}
int blockstore_impl_t::pad_journal_read(std::vector<copy_buffer_t> & rv, copy_buffer_t & cp,
// FIXME Passing dirty_entry& would be nicer
uint64_t dirty_offset, uint64_t dirty_end, uint64_t dirty_loc, uint8_t *csum_ptr, int *dyn_data,
@@ -598,11 +658,15 @@ bool blockstore_impl_t::fulfill_clean_read(blockstore_op_t *read_op, uint64_t &
{
auto & rv = PRIV(read_op)->read_vec;
int req = fill_partial_checksum_blocks(rv, fulfilled, clean_entry_bitmap, dyn_data, from_journal,
(uint8_t*)read_op->buf, read_op->offset, read_op->offset+read_op->len);
(uint8_t*)read_op->buf, read_op->offset, read_op->offset+read_op->len, item_start, item_end);
if (!inmemory_meta && !from_journal && req > 0)
{
// Read checksums from disk
uint8_t *csum_buf = read_clean_meta_block(read_op, clean_loc, rv.size()-req);
if (!csum_buf)
{
return false;
}
for (int i = req; i > 0; i--)
{
rv[rv.size()-i].csum_buf = csum_buf;
@@ -615,7 +679,7 @@ bool blockstore_impl_t::fulfill_clean_read(blockstore_op_t *read_op, uint64_t &
return false;
}
}
PRIV(read_op)->clean_block_used = req > 0;
PRIV(read_op)->clean_loc_used = req > 0 ? UINT64_MAX : 0;
}
else if (from_journal)
{
@@ -665,6 +729,10 @@ bool blockstore_impl_t::fulfill_clean_read(blockstore_op_t *read_op, uint64_t &
{
// Read checksums from disk
csum_buf = read_clean_meta_block(read_op, clean_loc, PRIV(read_op)->read_vec.size());
if (!csum_buf)
{
return false;
}
csum_done = true;
}
uint8_t *csum = !dsk.csum_block_size ? 0 : (csum_buf + 2*dsk.clean_entry_bitmap_size + bmp_start*(dsk.data_csum_type & 0xFF));
@@ -679,13 +747,13 @@ bool blockstore_impl_t::fulfill_clean_read(blockstore_op_t *read_op, uint64_t &
}
}
// Increment reference counter if clean data is being read from the disk
if (PRIV(read_op)->clean_block_used)
if (PRIV(read_op)->clean_loc_used == UINT64_MAX)
{
auto & uo = used_clean_objects[clean_loc];
uo.refs++;
if (dsk.csum_block_size && flusher->is_mutated(clean_loc))
uo.was_changed = true;
PRIV(read_op)->clean_block_used = clean_loc;
PRIV(read_op)->clean_loc_used = clean_loc + 1;
}
return true;
}
@@ -725,12 +793,18 @@ bool blockstore_impl_t::verify_padded_checksums(uint8_t *clean_entry_bitmap, uin
while (pos < iov[i].iov_len)
{
uint32_t start = pos;
uint8_t bit = (clean_entry_bitmap[bmp_pos >> 3] >> (bmp_pos & 0x7)) & 1;
while (pos < iov[i].iov_len && ((clean_entry_bitmap[bmp_pos >> 3] >> (bmp_pos & 0x7)) & 1) == bit)
uint8_t bit = 1;
if (clean_entry_bitmap)
{
pos += dsk.bitmap_granularity;
bmp_pos++;
bit = (clean_entry_bitmap[bmp_pos >> 3] >> (bmp_pos & 0x7)) & 1;
while (pos < iov[i].iov_len && ((clean_entry_bitmap[bmp_pos >> 3] >> (bmp_pos & 0x7)) & 1) == bit)
{
pos += dsk.bitmap_granularity;
bmp_pos++;
}
}
else
pos = iov[i].iov_len;
uint32_t len = pos-start;
auto buf = (uint8_t*)iov[i].iov_base+start;
while (block_done+len >= dsk.csum_block_size)
@@ -807,7 +881,7 @@ bool blockstore_impl_t::verify_clean_padded_checksums(blockstore_op_t *op, uint6
{
uint32_t offset = clean_loc % dsk.data_block_size;
if (from_journal)
return verify_padded_checksums(dyn_data, dyn_data + dsk.clean_entry_bitmap_size, offset, iov, n_iov, bad_block_cb);
return verify_padded_checksums(NULL, dyn_data, offset, iov, n_iov, bad_block_cb);
clean_loc = (clean_loc / dsk.data_block_size) * dsk.data_block_size;
if (!dyn_data)
{
@@ -835,7 +909,7 @@ void blockstore_impl_t::handle_read_event(ring_data_t *data, blockstore_op_t *op
void *meta_block = NULL;
if (dsk.csum_block_size > dsk.bitmap_granularity)
{
for (int i = rv.size()-1; i >= 0 && (rv[i].copy_flags & COPY_BUF_CSUM_FILL); i--)
for (int i = 0; i < rv.size(); i++)
{
if (rv[i].copy_flags & COPY_BUF_META_BLOCK)
{
@@ -845,8 +919,41 @@ void blockstore_impl_t::handle_read_event(ring_data_t *data, blockstore_op_t *op
rv[i].buf = NULL;
continue;
}
struct iovec *iov = (struct iovec*)((uint8_t*)rv[i].buf + (rv[i].len & 0xFFFFFFFF));
int n_iov = rv[i].len >> 32;
if (rv[i].copy_flags & COPY_BUF_ZERO)
{
// Zero read
continue;
}
if (rv[i].copy_flags & COPY_BUF_COALESCED)
{
// Sub-block shared with another read. Skip
continue;
}
if ((rv[i].copy_flags & COPY_BUF_JOURNAL) && journal.inmemory)
{
// Do not check journal checksums in-memory
continue;
}
iovec single_iov = {};
iovec *iov = NULL;
int n_iov = 0;
if (rv[i].copy_flags & COPY_BUF_CSUM_FILL)
{
// Padded, buffer list passed using a 'creepy way'
iov = (struct iovec*)((uint8_t*)rv[i].buf + (rv[i].len & 0xFFFFFFFF));
n_iov = rv[i].len >> 32;
}
else
{
// Not padded, buffer is fully within the input buffer
assert(op->buf);
assert(rv[i].csum_buf);
iov = &single_iov;
n_iov = 1;
assert(rv[i].offset >= op->offset);
assert(rv[i].offset + rv[i].len <= op->offset + op->len);
single_iov = { .iov_base = op->buf + rv[i].offset - op->offset, .iov_len = rv[i].len };
}
bool ok = true;
if (rv[i].copy_flags & COPY_BUF_JOURNAL)
{
@@ -944,23 +1051,7 @@ void blockstore_impl_t::handle_read_event(ring_data_t *data, blockstore_op_t *op
meta_block = NULL;
}
}
if (PRIV(op)->clean_block_used)
{
// Release clean data block
auto uo_it = used_clean_objects.find(PRIV(op)->clean_block_used);
if (uo_it != used_clean_objects.end())
{
uo_it->second.refs--;
if (uo_it->second.refs <= 0)
{
if (uo_it->second.was_freed)
{
data_alloc->set(PRIV(op)->clean_block_used, false);
}
used_clean_objects.erase(uo_it);
}
}
}
release_clean(op);
if (!journal.inmemory)
{
// Release journal sector usage
+2 -2
View File
@@ -491,7 +491,7 @@ void blockstore_impl_t::mark_stable(obj_ver_id v, bool forget_dirty)
if (!exists)
{
uint64_t space_id = dirty_it->first.oid.inode;
if (no_inode_stats[dirty_it->first.oid.inode >> (64-POOL_ID_BITS)])
if (no_inode_stats.find(dirty_it->first.oid.inode >> (64-POOL_ID_BITS)) != no_inode_stats.end())
space_id = space_id & ~(((uint64_t)1 << (64-POOL_ID_BITS)) - 1);
inode_space_stats[space_id] += dsk.data_block_size;
used_blocks++;
@@ -501,7 +501,7 @@ void blockstore_impl_t::mark_stable(obj_ver_id v, bool forget_dirty)
else if (IS_DELETE(dirty_it->second.state))
{
uint64_t space_id = dirty_it->first.oid.inode;
if (no_inode_stats[dirty_it->first.oid.inode >> (64-POOL_ID_BITS)])
if (no_inode_stats.find(dirty_it->first.oid.inode >> (64-POOL_ID_BITS)) != no_inode_stats.end())
space_id = space_id & ~(((uint64_t)1 << (64-POOL_ID_BITS)) - 1);
auto & sp = inode_space_stats[space_id];
if (sp > dsk.data_block_size)
+37 -11
View File
@@ -3,6 +3,22 @@ cmake_minimum_required(VERSION 2.8...3.30)
project(vitastor)
# libvitastor_common.a
add_library(vitastor_common STATIC
etcd_state_client.cpp
msgr_stop.cpp
msgr_op.cpp
../../json11/json11.cpp
osd_ops.cpp
pg_states.cpp
../util/allocator.cpp
../util/addr_util.cpp
../util/timerfd_manager.cpp
../util/str_util.cpp
../util/json_util.cpp
)
target_compile_options(vitastor_common PUBLIC -fPIC)
# libvitastor_net.a
set(MSGR_RDMA "")
if (IBVERBS_LIBRARIES)
set(MSGR_RDMA "msgr_rdma.cpp")
@@ -11,24 +27,32 @@ set(MSGR_RDMACM "")
if (RDMACM_LIBRARIES)
set(MSGR_RDMACM "msgr_rdmacm.cpp")
endif (RDMACM_LIBRARIES)
add_library(vitastor_common STATIC
../util/epoll_manager.cpp etcd_state_client.cpp messenger.cpp msgr_iothread.cpp ../util/addr_util.cpp
msgr_stop.cpp msgr_op.cpp msgr_send.cpp msgr_receive.cpp ../util/ringloop.cpp ../../json11/json11.cpp
http_client.cpp osd_ops.cpp pg_states.cpp ../util/timerfd_manager.cpp ../util/str_util.cpp ../util/json_util.cpp ${MSGR_RDMA} ${MSGR_RDMACM}
add_library(vitastor_net STATIC
../util/epoll_manager.cpp
etcd_state_client_http.cpp
messenger.cpp
msgr_iothread.cpp
msgr_send.cpp
msgr_receive.cpp
../util/ringloop.cpp
http_client.cpp
${MSGR_RDMA}
${MSGR_RDMACM}
)
target_link_libraries(vitastor_common pthread)
target_compile_options(vitastor_common PUBLIC -fPIC)
target_link_libraries(vitastor_net pthread vitastor_common)
target_compile_options(vitastor_net PUBLIC -fPIC)
# libvitastor_client.so
add_library(vitastor_client SHARED
cluster_client.cpp
cluster_client_real.cpp
cluster_client_list.cpp
cluster_client_wb.cpp
vitastor_c.cpp
)
set_target_properties(vitastor_client PROPERTIES PUBLIC_HEADER "client/vitastor_c.h")
target_link_libraries(vitastor_client
vitastor_common
vitastor_net
vitastor_cli
${LIBURING_LIBRARIES}
${IBVERBS_LIBRARIES}
@@ -95,11 +119,13 @@ endif (${WITH_QEMU})
add_executable(test_cluster_client
EXCLUDE_FROM_ALL
../test/test_cluster_client.cpp
pg_states.cpp osd_ops.cpp cluster_client.cpp cluster_client_list.cpp cluster_client_wb.cpp msgr_op.cpp ../test/mock/messenger.cpp msgr_stop.cpp
etcd_state_client.cpp ../util/timerfd_manager.cpp ../util/addr_util.cpp ../util/str_util.cpp ../util/json_util.cpp ../../json11/json11.cpp
cluster_client.cpp
cluster_client_list.cpp
cluster_client_wb.cpp
../test/mock/messenger.cpp
etcd_state_client_mock.cpp
)
target_link_libraries(test_cluster_client ${LIBURING_LIBRARIES})
target_compile_definitions(test_cluster_client PUBLIC -D__MOCK__)
target_link_libraries(test_cluster_client vitastor_common ${LIBURING_LIBRARIES})
target_include_directories(test_cluster_client BEFORE PUBLIC ${CMAKE_SOURCE_DIR}/src/test/mock)
add_dependencies(build_tests test_cluster_client)
add_test(NAME test_cluster_client COMMAND test_cluster_client)
+51 -51
View File
@@ -11,7 +11,7 @@
#define TRY_SEND_CONNECTING 1
#define TRY_SEND_OK 2
cluster_client_t::cluster_client_t(ring_loop_t *ringloop, timerfd_manager_t *tfd, json11::Json config)
cluster_client_t::cluster_client_t(ring_loop_t *ringloop, timerfd_manager_t *tfd, json11::Json config, std::unique_ptr<etcd_state_client_t> st_cli_ptr)
{
wb = new writeback_cache_t();
@@ -53,23 +53,23 @@ cluster_client_t::cluster_client_t(ring_loop_t *ringloop, timerfd_manager_t *tfd
};
msgr.parse_config(config);
st_cli.tfd = tfd;
st_cli.on_load_config_hook = [this](json11::Json::object & cfg) { on_load_config_hook(cfg); };
st_cli.on_change_osd_state_hook = [this](uint64_t peer_osd) { on_change_osd_state_hook(peer_osd); };
st_cli.on_change_pool_config_hook = [this]() { on_change_pool_config_hook(); };
st_cli.on_change_pg_config_hook = [this]() { on_change_pool_config_hook(); };
st_cli.on_change_pg_state_hook = [this](pool_id_t pool_id, pg_num_t pg_num, osd_num_t prev_primary) { on_change_pg_state_hook(pool_id, pg_num, prev_primary); };
st_cli.on_change_node_placement_hook = [this]() { on_change_node_placement_hook(); };
st_cli.on_load_pgs_hook = [this](bool success) { on_load_pgs_hook(success); };
st_cli.on_reload_hook = [this]() { st_cli.load_global_config(); };
st_cli = std::move(st_cli_ptr);
st_cli->on_load_config_hook = [this](json11::Json::object & cfg) { on_load_config_hook(cfg); };
st_cli->on_change_osd_state_hook = [this](uint64_t peer_osd) { on_change_osd_state_hook(peer_osd); };
st_cli->on_change_pool_config_hook = [this]() { on_change_pool_config_hook(); };
st_cli->on_change_pg_config_hook = [this]() { on_change_pool_config_hook(); };
st_cli->on_change_pg_state_hook = [this](pool_id_t pool_id, pg_num_t pg_num, osd_num_t prev_primary) { on_change_pg_state_hook(pool_id, pg_num, prev_primary); };
st_cli->on_change_node_placement_hook = [this]() { on_change_node_placement_hook(); };
st_cli->on_load_pgs_hook = [this](bool success) { on_load_pgs_hook(success); };
st_cli->on_reload_hook = [this]() { this->st_cli->load_global_config(); };
st_cli.parse_config(config);
st_cli.infinite_start = false;
st_cli->parse_config(config);
st_cli->infinite_start = false;
if (!config["client_infinite_start"].is_null())
{
st_cli.infinite_start = config["client_infinite_start"].bool_value();
st_cli->infinite_start = config["client_infinite_start"].bool_value();
}
st_cli.load_global_config();
st_cli->load_global_config();
scrap_buffer_size = SCRAP_BUFFER_SIZE;
scrap_buffer = malloc_or_die(scrap_buffer_size);
@@ -469,7 +469,7 @@ void cluster_client_t::on_load_config_hook(json11::Json::object & etcd_global_co
auto etcd_report_interval = config["etcd_report_interval"].uint64_value();
if (!etcd_report_interval)
etcd_report_interval = 5;
client_wait_up_timeout = 1+etcd_report_interval+(st_cli.max_etcd_attempts*(2*st_cli.etcd_quick_timeout)+999)/1000;
client_wait_up_timeout = 1+etcd_report_interval+(st_cli->max_etcd_attempts*(2*st_cli->etcd_quick_timeout)+999)/1000;
}
// log_level
log_level = config["log_level"].uint64_value();
@@ -482,8 +482,8 @@ void cluster_client_t::on_load_config_hook(json11::Json::object & etcd_global_co
client_hostname = new_hostname;
}
msgr.parse_config(config);
st_cli.parse_config(config);
st_cli.load_pgs();
st_cli->parse_config(config);
st_cli->load_pgs();
}
osd_num_t cluster_client_t::select_random_osd(const std::vector<osd_num_t> & osds)
@@ -492,7 +492,7 @@ osd_num_t cluster_client_t::select_random_osd(const std::vector<osd_num_t> & osd
int alive_count = 0;
for (auto & osd_num: osds)
{
if (!st_cli.peer_states[osd_num].is_null())
if (!st_cli->peer_states[osd_num].is_null())
alive_set[alive_count++] = osd_num;
}
if (!alive_count)
@@ -509,7 +509,7 @@ osd_num_t cluster_client_t::select_nearest_osd(const std::vector<osd_num_t> & os
while (self_tree_metrics.find(cur_id) == self_tree_metrics.end())
{
self_tree_metrics[cur_id] = metric++;
json11::Json cur_placement = st_cli.node_placement[cur_id];
json11::Json cur_placement = st_cli->node_placement[cur_id];
cur_id = cur_placement["parent"].string_value();
}
if (cur_id != "")
@@ -529,7 +529,7 @@ osd_num_t cluster_client_t::select_nearest_osd(const std::vector<osd_num_t> & os
}
else
{
auto & peer_state = st_cli.peer_states[osd_num];
auto & peer_state = st_cli->peer_states[osd_num];
if (!peer_state.is_null())
{
metric = self_tree_metrics[""];
@@ -539,7 +539,7 @@ osd_num_t cluster_client_t::select_nearest_osd(const std::vector<osd_num_t> & os
while (seen.find(cur_id) == seen.end())
{
seen.insert(cur_id);
json11::Json cur_placement = st_cli.node_placement[cur_id];
json11::Json cur_placement = st_cli->node_placement[cur_id];
std::string cur_parent = cur_placement["parent"].string_value();
cur_id = (!first || cur_parent != "" ? cur_parent : peer_state["host"].string_value());
first = false;
@@ -564,7 +564,7 @@ osd_num_t cluster_client_t::select_nearest_osd(const std::vector<osd_num_t> & os
void cluster_client_t::on_load_pgs_hook(bool success)
{
for (auto & pool_item: st_cli.pool_config)
for (auto & pool_item: st_cli->pool_config)
{
pg_counts[pool_item.first] = pool_item.second.real_pg_count;
}
@@ -584,7 +584,7 @@ void cluster_client_t::on_load_pgs_hook(bool success)
void cluster_client_t::on_change_pool_config_hook()
{
for (auto & pool_item: st_cli.pool_config)
for (auto & pool_item: st_cli->pool_config)
{
if (pg_counts[pool_item.first] != pool_item.second.real_pg_count)
{
@@ -612,7 +612,7 @@ void cluster_client_t::on_change_pool_config_hook()
void cluster_client_t::on_change_pg_state_hook(pool_id_t pool_id, pg_num_t pg_num, osd_num_t prev_primary)
{
auto & pg_cfg = st_cli.pool_config[pool_id].pg_config[pg_num];
auto & pg_cfg = st_cli->pool_config[pool_id].pg_config[pg_num];
if (pg_cfg.cur_primary != prev_primary)
{
// Repeat this PG operations because an OSD which stopped being primary may not fsync operations
@@ -630,8 +630,8 @@ bool cluster_client_t::get_immediate_commit(uint64_t inode)
pool_id_t pool_id = INODE_POOL(inode);
if (!pool_id)
return true;
auto pool_it = st_cli.pool_config.find(pool_id);
if (pool_it == st_cli.pool_config.end())
auto pool_it = st_cli->pool_config.find(pool_id);
if (pool_it == st_cli->pool_config.end())
return true;
return pool_it->second.immediate_commit == IMMEDIATE_ALL;
}
@@ -641,7 +641,7 @@ void cluster_client_t::on_change_osd_state_hook(uint64_t peer_osd)
osd_tree_metrics.erase(peer_osd);
if (msgr.wanted_peers.find(peer_osd) != msgr.wanted_peers.end())
{
msgr.connect_peer(peer_osd, st_cli.peer_states[peer_osd]);
msgr.connect_peer(peer_osd, st_cli->peer_states[peer_osd]);
continue_lists();
}
}
@@ -936,8 +936,8 @@ bool cluster_client_t::check_rw(cluster_op_t *op)
cb(op);
return false;
}
auto pool_it = st_cli.pool_config.find(pool_id);
if (pool_it == st_cli.pool_config.end() || pool_it->second.real_pg_count == 0)
auto pool_it = st_cli->pool_config.find(pool_id);
if (pool_it == st_cli->pool_config.end() || pool_it->second.real_pg_count == 0)
{
// Pools are loaded, but this one is unknown
op->retval = -EINVAL;
@@ -960,8 +960,8 @@ bool cluster_client_t::check_rw(cluster_op_t *op)
}
if ((op->opcode == OSD_OP_WRITE || op->opcode == OSD_OP_DELETE) && !(op->flags & OSD_OP_IGNORE_READONLY))
{
auto ino_it = st_cli.inode_config.find(op->inode);
if (ino_it != st_cli.inode_config.end() && ino_it->second.readonly)
auto ino_it = st_cli->inode_config.find(op->inode);
if (ino_it != st_cli->inode_config.end() && ino_it->second.readonly)
{
op->retval = -EROFS;
auto cb = std::move(op->callback);
@@ -972,15 +972,15 @@ bool cluster_client_t::check_rw(cluster_op_t *op)
op->deoptimise_snapshot = false;
if (enable_writeback && (op->opcode == OSD_OP_READ || op->opcode == OSD_OP_READ_BITMAP || op->opcode == OSD_OP_READ_CHAIN_BITMAP))
{
auto ino_it = st_cli.inode_config.find(op->inode);
if (ino_it != st_cli.inode_config.end())
auto ino_it = st_cli->inode_config.find(op->inode);
if (ino_it != st_cli->inode_config.end())
{
int chain_size = 0;
while (ino_it != st_cli.inode_config.end() && ino_it->second.parent_id)
while (ino_it != st_cli->inode_config.end() && ino_it->second.parent_id)
{
// Check for loops - FIXME check it in etcd_state_client
if (ino_it->second.parent_id == op->inode ||
chain_size > st_cli.inode_config.size())
chain_size > st_cli->inode_config.size())
{
op->retval = -EINVAL;
auto cb = std::move(op->callback);
@@ -995,7 +995,7 @@ bool cluster_client_t::check_rw(cluster_op_t *op)
break;
}
chain_size++;
ino_it = st_cli.inode_config.find(ino_it->second.parent_id);
ino_it = st_cli->inode_config.find(ino_it->second.parent_id);
}
}
}
@@ -1014,7 +1014,7 @@ void cluster_client_t::execute_raw(osd_num_t osd_num, osd_op_t *op)
else
{
if (msgr.wanted_peers.find(osd_num) == msgr.wanted_peers.end())
msgr.connect_peer(osd_num, st_cli.peer_states[osd_num]);
msgr.connect_peer(osd_num, st_cli->peer_states[osd_num]);
raw_ops.emplace(osd_num, op);
}
}
@@ -1129,25 +1129,25 @@ resume_2:
if (op->opcode == OSD_OP_READ || op->opcode == OSD_OP_READ_CHAIN_BITMAP)
{
// Check parent inode
auto ino_it = st_cli.inode_config.find(op->cur_inode);
auto ino_it = st_cli->inode_config.find(op->cur_inode);
// Skip parents from the same pool
int skipped = 0;
while (!op->deoptimise_snapshot &&
ino_it != st_cli.inode_config.end() && ino_it->second.parent_id &&
ino_it != st_cli->inode_config.end() && ino_it->second.parent_id &&
INODE_POOL(ino_it->second.parent_id) == INODE_POOL(op->cur_inode))
{
// Check for loops - FIXME check it in etcd_state_client
if (ino_it->second.parent_id == op->inode ||
skipped > st_cli.inode_config.size())
skipped > st_cli->inode_config.size())
{
op->retval = -EINVAL;
erase_op(op);
return 1;
}
skipped++;
ino_it = st_cli.inode_config.find(ino_it->second.parent_id);
ino_it = st_cli->inode_config.find(ino_it->second.parent_id);
}
if (ino_it != st_cli.inode_config.end() &&
if (ino_it != st_cli->inode_config.end() &&
ino_it->second.parent_id &&
ino_it->second.parent_id != op->inode)
{
@@ -1161,7 +1161,7 @@ resume_2:
op->retval = op->len;
if (op->opcode == OSD_OP_READ_BITMAP || op->opcode == OSD_OP_READ_CHAIN_BITMAP)
{
auto & pool_cfg = st_cli.pool_config.at(INODE_POOL(op->inode));
auto & pool_cfg = st_cli->pool_config.at(INODE_POOL(op->inode));
op->retval = op->len / pool_cfg.bitmap_granularity;
}
if (op->flush_id)
@@ -1247,7 +1247,7 @@ void cluster_client_t::slice_rw(cluster_op_t *op)
{
// Slice the request into individual object stripe requests
// Primary OSDs still operate individual stripes, but their size is multiplied by PG minsize in case of EC
auto & pool_cfg = st_cli.pool_config.at(INODE_POOL(op->cur_inode));
auto & pool_cfg = st_cli->pool_config.at(INODE_POOL(op->cur_inode));
uint32_t pg_data_size = (pool_cfg.scheme == POOL_SCHEME_REPLICATED ? 1 : pool_cfg.pg_size-pool_cfg.parity_chunks);
uint64_t pg_block_size = pool_cfg.data_block_size * pg_data_size;
uint64_t first_stripe = (op->offset / pg_block_size) * pg_block_size;
@@ -1346,7 +1346,7 @@ bool cluster_client_t::affects_pg(uint64_t inode, uint64_t offset, uint64_t len,
{
return false;
}
auto & pool_cfg = st_cli.pool_config.at(INODE_POOL(inode));
auto & pool_cfg = st_cli->pool_config.at(INODE_POOL(inode));
uint32_t pg_data_size = (pool_cfg.scheme == POOL_SCHEME_REPLICATED ? 1 : pool_cfg.pg_size-pool_cfg.parity_chunks);
uint64_t pg_block_size = pool_cfg.data_block_size * pg_data_size;
uint64_t first_stripe = (offset / pg_block_size) * pg_block_size;
@@ -1365,7 +1365,7 @@ bool cluster_client_t::affects_pg(uint64_t inode, uint64_t offset, uint64_t len,
bool cluster_client_t::affects_osd(uint64_t inode, uint64_t offset, uint64_t len, osd_num_t osd)
{
auto & pool_cfg = st_cli.pool_config.at(INODE_POOL(inode));
auto & pool_cfg = st_cli->pool_config.at(INODE_POOL(inode));
uint32_t pg_data_size = (pool_cfg.scheme == POOL_SCHEME_REPLICATED ? 1 : pool_cfg.pg_size-pool_cfg.parity_chunks);
uint64_t pg_block_size = pool_cfg.data_block_size * pg_data_size;
uint64_t first_stripe = (offset / pg_block_size) * pg_block_size;
@@ -1389,7 +1389,7 @@ int cluster_client_t::try_send(cluster_op_t *op, int i, std::function<void(osd_o
init_msgr();
}
auto part = &op->parts[i];
auto & pool_cfg = st_cli.pool_config.at(INODE_POOL(op->cur_inode));
auto & pool_cfg = st_cli->pool_config.at(INODE_POOL(op->cur_inode));
auto pg_it = pool_cfg.pg_config.find(part->pg_num);
if (pg_it != pool_cfg.pg_config.end() &&
!pg_it->second.pause && pg_it->second.cur_primary &&
@@ -1420,8 +1420,8 @@ int cluster_client_t::try_send(cluster_op_t *op, int i, std::function<void(osd_o
uint64_t meta_rev = 0;
if (op->opcode != OSD_OP_READ_BITMAP && op->opcode != OSD_OP_DELETE && !op->deoptimise_snapshot)
{
auto ino_it = st_cli.inode_config.find(op->cur_inode);
if (ino_it != st_cli.inode_config.end())
auto ino_it = st_cli->inode_config.find(op->cur_inode);
if (ino_it != st_cli->inode_config.end())
meta_rev = ino_it->second.mod_revision;
}
part->op = (osd_op_t){
@@ -1453,7 +1453,7 @@ int cluster_client_t::try_send(cluster_op_t *op, int i, std::function<void(osd_o
}
else if (msgr.wanted_peers.find(primary_osd) == msgr.wanted_peers.end())
{
msgr.connect_peer(primary_osd, st_cli.peer_states[primary_osd]);
msgr.connect_peer(primary_osd, st_cli->peer_states[primary_osd]);
return TRY_SEND_CONNECTING;
}
}
@@ -1648,7 +1648,7 @@ void cluster_client_t::handle_op_part(cluster_op_part_t *part)
void cluster_client_t::copy_part_bitmap(cluster_op_t *op, cluster_op_part_t *part)
{
// Copy (OR) bitmap
auto & pool_cfg = st_cli.pool_config.at(INODE_POOL(op->cur_inode));
auto & pool_cfg = st_cli->pool_config.at(INODE_POOL(op->cur_inode));
uint32_t pg_block_size = pool_cfg.data_block_size * (
pool_cfg.scheme == POOL_SCHEME_REPLICATED ? 1 : pool_cfg.pg_size-pool_cfg.parity_chunks
);
+4 -3
View File
@@ -4,7 +4,7 @@
#pragma once
#include "messenger.h"
#include "etcd_state_client.h"
#include "etcd_state_client_http.h"
#define DEFAULT_CLIENT_MAX_DIRTY_BYTES 32*1024*1024
#define DEFAULT_CLIENT_MAX_DIRTY_OPS 1024
@@ -131,7 +131,7 @@ class __attribute__((visibility("default"))) cluster_client_t
bool msgr_initialized = false;
public:
etcd_state_client_t st_cli;
std::unique_ptr<etcd_state_client_t> st_cli;
osd_messenger_t msgr;
void init_msgr();
@@ -139,7 +139,8 @@ public:
json11::Json::object cli_config, file_config, etcd_global_config;
json11::Json::object config;
cluster_client_t(ring_loop_t *ringloop, timerfd_manager_t *tfd, json11::Json config);
static cluster_client_t* create(ring_loop_t *ringloop, timerfd_manager_t *tfd, json11::Json config);
cluster_client_t(ring_loop_t *ringloop, timerfd_manager_t *tfd, json11::Json config, std::unique_ptr<etcd_state_client_t> st_cli);
~cluster_client_t();
void execute(cluster_op_t *op);
void execute_raw(osd_num_t osd_num, osd_op_t *op);
+10 -10
View File
@@ -63,14 +63,14 @@ void cluster_client_t::list_inode(inode_t inode, uint64_t min_offset, uint64_t m
{
init_msgr();
pool_id_t pool_id = INODE_POOL(inode);
if (!pool_id || st_cli.pool_config.find(pool_id) == st_cli.pool_config.end())
if (!pool_id || st_cli->pool_config.find(pool_id) == st_cli->pool_config.end())
{
if (log_level > 0)
fprintf(stderr, "Pool %u does not exist\n", pool_id);
pg_callback(-EINVAL, 0, 0, std::set<object_id>());
return;
}
auto pg_stripe_size = st_cli.pool_config.at(pool_id).pg_stripe_size;
auto pg_stripe_size = st_cli->pool_config.at(pool_id).pg_stripe_size;
if (min_offset)
min_offset = (min_offset/pg_stripe_size) * pg_stripe_size;
inode_list_t *lst = new inode_list_t();
@@ -110,13 +110,13 @@ bool cluster_client_t::continue_listing(inode_list_t *lst)
bool cluster_client_t::restart_listing(inode_list_t* lst)
{
auto pool_it = st_cli.pool_config.find(lst->pool_id);
auto pool_it = st_cli->pool_config.find(lst->pool_id);
// We want listing to be consistent. To achieve it we should:
// 1) retry listing of each PG if its state changes
// 2) abort listing if PG count changes during listing
// 3) ideally, only talk to the primary OSD - this will be done separately
// So first we add all PGs without checking their state
if (pool_it == st_cli.pool_config.end() ||
if (pool_it == st_cli->pool_config.end() ||
lst->real_pg_count != pool_it->second.real_pg_count)
{
for (auto pg: lst->pgs)
@@ -136,7 +136,7 @@ bool cluster_client_t::restart_listing(inode_list_t* lst)
fprintf(stderr, "PG count in pool %u changed during listing\n", lst->pool_id);
}
lst->pgs.clear();
if (pool_it == st_cli.pool_config.end())
if (pool_it == st_cli->pool_config.end())
{
// Unknown pool
lst->callback(-EINVAL, 0, 0, std::set<object_id>());
@@ -248,7 +248,7 @@ void cluster_client_t::set_list_retry_timeout(int ms, timespec new_time)
int cluster_client_t::start_pg_listing(inode_list_pg_t *pg)
{
auto & pool_cfg = st_cli.pool_config.at(pg->lst->pool_id);
auto & pool_cfg = st_cli->pool_config.at(pg->lst->pool_id);
auto pg_it = pool_cfg.pg_config.find(pg->pg_num);
assert(pg->lst->real_pg_count == pool_cfg.real_pg_count);
if (pg_it == pool_cfg.pg_config.end() ||
@@ -277,7 +277,7 @@ int cluster_client_t::start_pg_listing(inode_list_pg_t *pg)
for (auto peer_it = all_peers.begin(); peer_it != all_peers.end(); )
{
if (*peer_it != pg_it->second.cur_primary &&
st_cli.peer_states[*peer_it].is_null())
st_cli->peer_states[*peer_it].is_null())
{
pg->inactive_osds.push_back(*peer_it);
all_peers.erase(peer_it++);
@@ -298,11 +298,11 @@ int cluster_client_t::start_pg_listing(inode_list_pg_t *pg)
if (msgr.osd_peers.find(peer_osd) == msgr.osd_peers.end())
{
// Initiate connection
if (st_cli.peer_states[peer_osd].is_null())
if (st_cli->peer_states[peer_osd].is_null())
{
return LIST_PG_WAIT_ACTIVE;
}
msgr.connect_peer(peer_osd, st_cli.peer_states[peer_osd]);
msgr.connect_peer(peer_osd, st_cli->peer_states[peer_osd]);
conn = false;
}
}
@@ -336,7 +336,7 @@ void cluster_client_t::send_list(inode_list_osd_t *cur_list)
if (!cur_list->pg->inflight_ops)
cur_list->pg->lst->inflight_pgs++;
cur_list->pg->inflight_ops++;
auto & pool_cfg = st_cli.pool_config[cur_list->pg->lst->pool_id];
auto & pool_cfg = st_cli->pool_config[cur_list->pg->lst->pool_id];
osd_op_t *op = new osd_op_t();
op->op_type = OSD_OP_OUT;
// Already checked that it exists above, but anyway
+11
View File
@@ -0,0 +1,11 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 or GNU GPL-2.0+ (see README.md for details)
#include "cluster_client.h"
#include "etcd_state_client_http.h"
cluster_client_t* cluster_client_t::create(ring_loop_t *ringloop, timerfd_manager_t *tfd, json11::Json config)
{
auto st_cli = new etcd_state_client_http_t(tfd);
return new cluster_client_t(ringloop, tfd, config, std::unique_ptr<etcd_state_client_t>(st_cli));
}
+2
View File
@@ -131,6 +131,7 @@ void writeback_cache_t::copy_write(cluster_op_t *op, int state, uint64_t new_flu
writeback_bytes -= op->len;
}
writeback_queue_size++;
writeback_queue.push_back({ op->inode, new_end });
}
break;
}
@@ -165,6 +166,7 @@ void writeback_cache_t::copy_write(cluster_op_t *op, int state, uint64_t new_flu
{
writeback_queue_size++;
}
writeback_queue.push_back({ op->inode, new_end });
}
auto new_dirty_it = dirty_buffers.emplace_hint(dirty_it, (object_id){
.inode = op->inode,
+20 -462
View File
@@ -1,13 +1,12 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 or GNU GPL-2.0+ (see README.md for details)
#include <assert.h>
#include "osd_ops.h"
#include "pg_states.h"
#include "etcd_state_client.h"
#ifndef __MOCK__
#include "addr_util.h"
#include "http_client.h"
#endif
#include "str_util.h"
etcd_state_client_t::~etcd_state_client_t()
@@ -17,28 +16,8 @@ etcd_state_client_t::~etcd_state_client_t()
delete watch;
}
watches.clear();
etcd_watches_initialised = -1;
#ifndef __MOCK__
stop_ws_keepalive();
if (etcd_watch_ws)
{
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
}
if (keepalive_client)
{
http_close(keepalive_client);
keepalive_client = NULL;
}
#endif
if (load_pgs_timer_id >= 0)
{
tfd->clear_timer(load_pgs_timer_id);
load_pgs_timer_id = -1;
}
}
#ifndef __MOCK__
etcd_kv_t etcd_state_client_t::parse_etcd_kv(const json11::Json & kv_json)
{
etcd_kv_t kv;
@@ -72,104 +51,6 @@ std::vector<std::string> etcd_state_client_t::get_addresses()
return addrs;
}
void etcd_state_client_t::etcd_call_oneshot(std::string etcd_address, std::string api, json11::Json payload,
int timeout, std::function<void(std::string, json11::Json)> callback)
{
std::string etcd_api_path;
int pos = etcd_address.find('/');
if (pos >= 0)
{
etcd_api_path = etcd_address.substr(pos);
etcd_address = etcd_address.substr(0, pos);
}
std::string req = payload.dump();
req = "POST "+etcd_api_path+api+" HTTP/1.1\r\n"
"Host: "+etcd_address+"\r\n"
"Content-Type: application/json\r\n"
"Content-Length: "+std::to_string(req.size())+"\r\n"
"Connection: close\r\n"
"\r\n"+req;
auto http_cli = http_init(tfd);
auto cb = [http_cli, callback](const http_response_t *response)
{
std::string err;
json11::Json data;
response->parse_json_response(err, data);
callback(err, data);
http_close(http_cli);
};
http_request(http_cli, etcd_address, req, { .timeout = timeout }, cb);
}
void etcd_state_client_t::etcd_call(std::string api, json11::Json payload, int timeout,
int retries, int interval, std::function<void(std::string, json11::Json)> callback)
{
if (!etcd_addresses.size() && !etcd_local.size())
{
fprintf(stderr, "etcd_address is missing in Vitastor configuration\n");
exit(1);
}
pick_next_etcd();
std::string etcd_address = selected_etcd_address;
std::string etcd_api_path;
int pos = etcd_address.find('/');
if (pos >= 0)
{
etcd_api_path = etcd_address.substr(pos);
etcd_address = etcd_address.substr(0, pos);
}
std::string req = payload.dump();
req = "POST "+etcd_api_path+api+" HTTP/1.1\r\n"
"Host: "+etcd_address+"\r\n"
"Content-Type: application/json\r\n"
"Content-Length: "+std::to_string(req.size())+"\r\n"
"Connection: keep-alive\r\n"
"Keep-Alive: timeout="+std::to_string(etcd_keepalive_timeout)+"\r\n"
"\r\n"+req;
retries--;
auto cb = [this, api, payload, timeout, retries, interval, callback,
cur_addr = selected_etcd_address](const http_response_t *response)
{
std::string err;
json11::Json data;
response->parse_json_response(err, data);
if (err != "")
{
if (cur_addr == selected_etcd_address)
selected_etcd_address = "";
if (retries > 0)
{
if (this->log_level > 0)
{
fprintf(
stderr, "Warning: etcd request failed: %s, retrying %d more times\n",
err.c_str(), retries
);
}
if (interval > 0)
{
// FIXME: Prevent destruction of etcd_state_client if timers or requests are active
tfd->set_timer(interval, false, [this, api, payload, timeout, retries, interval, callback](int)
{
etcd_call(api, payload, timeout, retries, interval, callback);
});
}
else
etcd_call(api, payload, timeout, retries, interval, callback);
}
else
callback(err, data);
}
else
callback(err, data);
};
if (!keepalive_client)
{
keepalive_client = http_init(tfd);
}
http_request(keepalive_client, etcd_address, req, { .timeout = timeout, .keepalive = true }, cb);
}
void etcd_state_client_t::add_etcd_url(std::string addr)
{
if (addr.length() > 0)
@@ -256,7 +137,6 @@ void etcd_state_client_t::parse_config(const json11::Json & config)
if (this->etcd_keepalive_timeout < 30)
this->etcd_keepalive_timeout = 30;
}
auto old_etcd_ws_keepalive_interval = this->etcd_ws_keepalive_interval;
this->etcd_ws_keepalive_interval = config["etcd_ws_keepalive_interval"].uint64_value();
if (this->etcd_ws_keepalive_interval <= 0)
{
@@ -282,294 +162,9 @@ void etcd_state_client_t::parse_config(const json11::Json & config)
{
this->etcd_min_reload_interval = 50;
}
if (this->etcd_ws_keepalive_interval != old_etcd_ws_keepalive_interval && ws_keepalive_timer >= 0)
{
#ifndef __MOCK__
stop_ws_keepalive();
start_ws_keepalive();
#endif
}
}
void etcd_state_client_t::pick_next_etcd()
{
if (selected_etcd_address != "")
return;
if (addresses_to_try.size() == 0)
{
// Prefer local etcd, if any
for (int i = 0; i < etcd_local.size(); i++)
addresses_to_try.push_back(etcd_local[i]);
std::vector<int> ns;
for (int i = 0; i < etcd_addresses.size(); i++)
ns.push_back(i);
if (!rand_initialized)
{
timespec tv;
clock_gettime(CLOCK_REALTIME, &tv);
srand48(tv.tv_sec*1000000000 + tv.tv_nsec);
rand_initialized = true;
}
while (ns.size())
{
int i = lrand48() % ns.size();
addresses_to_try.push_back(etcd_addresses[ns[i]]);
ns.erase(ns.begin()+i, ns.begin()+i+1);
}
}
selected_etcd_address = addresses_to_try[0];
addresses_to_try.erase(addresses_to_try.begin(), addresses_to_try.begin()+1);
}
void etcd_state_client_t::start_etcd_watcher()
{
if (!etcd_addresses.size() && !etcd_local.size())
{
fprintf(stderr, "etcd_address is missing in Vitastor configuration\n");
exit(1);
}
pick_next_etcd();
std::string etcd_address = selected_etcd_address;
std::string etcd_api_path;
int pos = etcd_address.find('/');
if (pos >= 0)
{
etcd_api_path = etcd_address.substr(pos);
etcd_address = etcd_address.substr(0, pos);
}
etcd_watches_initialised = 0;
ws_alive = 1;
if (etcd_watch_ws)
{
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
}
if (this->log_level > 1)
{
fprintf(stderr, "Trying to connect to etcd websocket at %s, watch from revision %ju/%ju/%ju\n", etcd_address.c_str(),
etcd_watch_revision_config, etcd_watch_revision_osd, etcd_watch_revision_pg);
}
etcd_watch_ws = open_websocket(tfd, etcd_address, etcd_api_path+"/watch", etcd_slow_timeout,
[this, cur_addr = selected_etcd_address](const http_response_t *msg)
{
if (msg->body.length())
{
ws_alive = 1;
std::string json_err;
json11::Json data = json11::Json::parse(msg->body, json_err);
if (json_err != "")
{
fprintf(stderr, "Bad JSON in etcd event: %s, ignoring event\n", json_err.c_str());
}
else
{
uint64_t watch_id = data["result"]["watch_id"].uint64_value();
if (data["result"]["created"].bool_value())
{
if (watch_id == ETCD_CONFIG_WATCH_ID ||
watch_id == ETCD_PG_STATE_WATCH_ID ||
watch_id == ETCD_OSD_STATE_WATCH_ID)
{
etcd_watches_initialised++;
}
if (etcd_watches_initialised == ETCD_TOTAL_WATCHES && this->log_level > 0)
{
fprintf(stderr, "Successfully subscribed to etcd at %s, revision %ju/%ju/%ju\n", cur_addr.c_str(),
etcd_watch_revision_config, etcd_watch_revision_osd, etcd_watch_revision_pg);
}
}
if (data["result"]["canceled"].bool_value())
{
// etcd watch canceled, maybe because the revision was compacted
if (data["result"]["compact_revision"].uint64_value())
{
// we may miss events if we proceed
// so we should restart from the beginning if we can
if (on_reload_hook != NULL)
{
// check to not trigger on_reload_hook multiple times
if (etcd_watch_ws != NULL)
{
fprintf(stderr, "Revisions before %ju were compacted by etcd, reloading state\n",
data["result"]["compact_revision"].uint64_value());
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
etcd_watch_revision_config = etcd_watch_revision_osd = etcd_watch_revision_pg = 0;
on_reload_hook();
}
return;
}
else
{
fprintf(stderr, "Revisions before %ju were compacted by etcd, exiting\n",
data["result"]["compact_revision"].uint64_value());
exit(1);
}
}
else
{
fprintf(stderr, "Watch canceled by etcd, reason: %s, exiting\n", data["result"]["cancel_reason"].string_value().c_str());
exit(1);
}
}
// Save revision only if it's present in the message - because sometimes etcd sends something without a header, like:
// {"error": {"grpc_code": 14, "http_code": 503, "http_status": "Service Unavailable", "message": "error reading from server: EOF"}}
// Also don't save revision from the initial created: true messages because they always contain the latest revision
if (etcd_watches_initialised == ETCD_TOTAL_WATCHES &&
!data["result"]["header"]["revision"].is_null() &&
!data["result"]["created"].bool_value())
{
// Restart watchers from the same revision number as in the last received message,
// not from the next one to protect against revision being split into multiple messages,
// even though etcd guarantees not to do that **within a single watcher** without fragment=true:
// https://etcd.io/docs/v3.5/learning/api_guarantees/#watch-apis
// Revision contents are ALWAYS split into separate messages for different watchers though!
// So generally we have to resume each watcher from its own revision...
// Progress messages may have watch_id=-1 if sent on behalf of multiple watchers though.
// And antietcd has an advanced semantic which merges the same revision for all watchers
// into one message and just omits watch_id.
// So we also have to handle the case where watch_id is -1 or not present (0).
auto watch_rev = data["result"]["header"]["revision"].uint64_value();
if (!watch_id || watch_id == UINT64_MAX)
etcd_watch_revision_config = etcd_watch_revision_osd = etcd_watch_revision_pg = watch_rev;
else if (watch_id == ETCD_CONFIG_WATCH_ID)
etcd_watch_revision_config = watch_rev;
else if (watch_id == ETCD_PG_STATE_WATCH_ID)
etcd_watch_revision_pg = watch_rev;
else if (watch_id == ETCD_OSD_STATE_WATCH_ID)
etcd_watch_revision_osd = watch_rev;
addresses_to_try.clear();
}
// First gather all changes into a hash to remove multiple overwrites
std::map<std::string, etcd_kv_t> changes;
for (auto & ev: data["result"]["events"].array_items())
{
auto kv = parse_etcd_kv(ev["kv"]);
if (kv.key != "")
{
changes[kv.key] = kv;
}
}
for (auto & kv: changes)
{
if (this->log_level > 3)
{
fprintf(stderr, "Incoming event: %s -> %s\n", kv.first.c_str(), kv.second.value.dump().c_str());
}
parse_state(kv.second);
}
// React to changes
if (on_change_hook != NULL)
{
on_change_hook(changes);
}
}
}
if (msg->eof)
{
fprintf(stderr, "Disconnected from etcd %s\n", cur_addr.c_str());
if (cur_addr == selected_etcd_address)
selected_etcd_address = "";
if (etcd_watch_ws)
{
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
}
if (etcd_watches_initialised == 0)
{
// Connection not established, retry in <etcd_quick_timeout>
tfd->set_timer(etcd_quick_timeout, false, [this](int)
{
start_etcd_watcher();
});
}
else if (etcd_watches_initialised > 0)
{
// Connection was live, retry immediately
etcd_watches_initialised = 0;
start_etcd_watcher();
}
}
});
http_post_message(etcd_watch_ws, WS_TEXT, json11::Json(json11::Json::object {
{ "create_request", json11::Json::object {
{ "key", base64_encode(etcd_prefix+"/config/") },
{ "range_end", base64_encode(etcd_prefix+"/config0") },
{ "start_revision", etcd_watch_revision_config },
{ "watch_id", ETCD_CONFIG_WATCH_ID },
{ "progress_notify", true },
} }
}).dump());
http_post_message(etcd_watch_ws, WS_TEXT, json11::Json(json11::Json::object {
{ "create_request", json11::Json::object {
{ "key", base64_encode(etcd_prefix+"/osd/state/") },
{ "range_end", base64_encode(etcd_prefix+"/osd/state0") },
{ "start_revision", etcd_watch_revision_osd },
{ "watch_id", ETCD_OSD_STATE_WATCH_ID },
{ "progress_notify", true },
} }
}).dump());
http_post_message(etcd_watch_ws, WS_TEXT, json11::Json(json11::Json::object {
{ "create_request", json11::Json::object {
{ "key", base64_encode(etcd_prefix+"/pg/") },
{ "range_end", base64_encode(etcd_prefix+"/pg0") },
{ "start_revision", etcd_watch_revision_pg },
{ "watch_id", ETCD_PG_STATE_WATCH_ID },
{ "progress_notify", true },
} }
}).dump());
// FIXME: Do not watch /pg/history/ at all in client code (not in OSD)
if (on_start_watcher_hook)
{
on_start_watcher_hook(etcd_watch_ws);
}
start_ws_keepalive();
}
void etcd_state_client_t::stop_ws_keepalive()
{
if (ws_keepalive_timer >= 0)
{
tfd->clear_timer(ws_keepalive_timer);
ws_keepalive_timer = -1;
}
}
void etcd_state_client_t::start_ws_keepalive()
{
if (ws_keepalive_timer < 0)
{
ws_keepalive_timer = tfd->set_timer(etcd_ws_keepalive_interval*1000, true, [this](int)
{
if (!etcd_watch_ws || etcd_watches_initialised < ETCD_TOTAL_WATCHES)
{
// Do nothing
}
else if (!ws_alive)
{
if (this->log_level > 0)
{
fprintf(stderr, "Websocket ping failed, disconnecting from etcd %s\n", selected_etcd_address.c_str());
}
if (etcd_watch_ws)
{
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
}
start_etcd_watcher();
}
else
{
ws_alive = 0;
http_post_message(etcd_watch_ws, WS_TEXT, json11::Json(json11::Json::object {
{ "progress_request", json11::Json::object { } }
}).dump());
}
});
}
}
void etcd_state_client_t::load_global_config()
void etcd_state_client_t::load_global_config(std::function<void(const std::string & error)> cb)
{
json11::Json::object req = { { "success", json11::Json::array {
json11::Json::object {
@@ -583,22 +178,12 @@ void etcd_state_client_t::load_global_config()
} }
},
} } };
etcd_txn(req, etcd_quick_timeout, max_etcd_attempts, 0, [this](std::string err, json11::Json data)
etcd_txn(req, etcd_quick_timeout, max_etcd_attempts, 0, [this, cb](std::string err, json11::Json data)
{
if (err != "")
{
fprintf(stderr, "Error reading configuration from etcd: %s\n", err.c_str());
if (infinite_start)
{
tfd->set_timer(etcd_slow_timeout, false, [this](int timer_id)
{
load_global_config();
});
}
else
{
exit(1);
}
cb(err);
return;
}
json11::Json config_kv = data["responses"][0]["response_range"]["kvs"][0];
@@ -629,28 +214,12 @@ void etcd_state_client_t::load_global_config()
parse_state(kv);
}
on_load_config_hook(global_config);
cb("");
});
}
void etcd_state_client_t::load_pgs()
void etcd_state_client_t::load_pgs(std::function<void(const std::string &)> cb)
{
timespec tv;
clock_gettime(CLOCK_REALTIME, &tv);
uint64_t ms_passed = (tv.tv_sec-etcd_last_reload.tv_sec)*1000 + (tv.tv_nsec-etcd_last_reload.tv_nsec)/1000000;
if (ms_passed < etcd_min_reload_interval)
{
if (load_pgs_timer_id < 0)
{
load_pgs_timer_id = tfd->set_timer(etcd_min_reload_interval+50-ms_passed, false, [this](int) { load_pgs(); });
}
return;
}
etcd_last_reload = tv;
if (load_pgs_timer_id >= 0)
{
tfd->clear_timer(load_pgs_timer_id);
load_pgs_timer_id = -1;
}
json11::Json::array txn = {
json11::Json::object {
{ "request_range", json11::Json::object {
@@ -698,16 +267,13 @@ void etcd_state_client_t::load_pgs()
{
req["compare"] = checks;
}
etcd_txn_slow(req, [this](std::string err, json11::Json data)
etcd_txn_slow(req, [this, cb](std::string err, json11::Json data)
{
if (err != "")
{
// Retry indefinitely
fprintf(stderr, "Error loading PGs from etcd: %s\n", err.c_str());
tfd->set_timer(etcd_slow_timeout, false, [this](int timer_id)
{
load_pgs();
});
cb(err);
return;
}
if (!data["succeeded"].bool_value())
@@ -735,24 +301,9 @@ void etcd_state_client_t::load_pgs()
}
clean_nonexistent_pgs();
on_load_pgs_hook(true);
start_etcd_watcher();
cb("");
});
}
#else
void etcd_state_client_t::parse_config(const json11::Json & config)
{
}
void etcd_state_client_t::load_global_config()
{
json11::Json::object global_config;
on_load_config_hook(global_config);
}
void etcd_state_client_t::load_pgs()
{
}
#endif
void etcd_state_client_t::reset_pg_exists()
{
@@ -865,7 +416,8 @@ void etcd_state_client_t::parse_state(const etcd_kv_t & kv)
if (pc.pg_size < 1 ||
pool_item.second["pg_size"].uint64_value() < 3 &&
(pc.scheme == POOL_SCHEME_XOR || pc.scheme == POOL_SCHEME_EC) ||
pool_item.second["pg_size"].uint64_value() > 256)
// limit is 64 because osd_peering_pg.cpp uses a 64-bit mask for has_roles
pool_item.second["pg_size"].uint64_value() > 64)
{
fprintf(stderr, "Pool %u has invalid pg_size, skipping pool\n", pool_id);
continue;
@@ -1205,8 +757,14 @@ void etcd_state_client_t::parse_state(const etcd_kv_t & kv)
else if (key.substr(0, etcd_prefix.length()+11) == etcd_prefix+"/osd/state/")
{
// <etcd_prefix>/osd/state/%d
osd_num_t peer_osd = std::stoull(key.substr(etcd_prefix.length()+11));
if (peer_osd > 0)
osd_num_t peer_osd = 0;
char null_byte = 0;
int scanned = sscanf(key.c_str() + etcd_prefix.length()+11, "%ju%c", &peer_osd, &null_byte);
if (scanned != 1 || !peer_osd)
{
fprintf(stderr, "Bad etcd key %s, ignoring\n", key.c_str());
}
else
{
if (value.is_object() && value["state"] == "up")
{
+14 -24
View File
@@ -103,15 +103,13 @@ protected:
std::vector<std::string> local_ips;
std::vector<std::string> etcd_addresses;
std::vector<std::string> etcd_local;
std::string selected_etcd_address;
std::vector<std::string> addresses_to_try;
std::vector<inode_watch_t*> watches;
std::set<osd_num_t> seen_peers;
bool new_pg_config = false;
int ws_keepalive_timer = -1;
int ws_alive = 0;
bool rand_initialized = false;
void add_etcd_url(std::string);
void pick_next_etcd();
void reset_pg_exists();
void clean_nonexistent_pgs();
public:
int etcd_keepalive_timeout = 30;
int etcd_ws_keepalive_interval = 5;
@@ -123,21 +121,15 @@ public:
uint64_t global_block_size = DEFAULT_BLOCK_SIZE;
uint32_t global_bitmap_granularity = DEFAULT_BITMAP_GRANULARITY;
uint32_t global_immediate_commit = IMMEDIATE_NONE;
std::string etcd_prefix;
int log_level = 0;
timerfd_manager_t *tfd = NULL;
http_co_t *etcd_watch_ws = NULL, *keepalive_client = NULL;
int etcd_watches_initialised = 0;
uint64_t etcd_watch_revision_config = 0;
uint64_t etcd_watch_revision_osd = 0;
uint64_t etcd_watch_revision_pg = 0;
timespec etcd_last_reload = {};
int load_pgs_timer_id = -1;
std::map<pool_id_t, pool_config_t> pool_config;
std::map<osd_num_t, json11::Json> peer_states;
std::set<osd_num_t> seen_peers;
std::map<inode_t, inode_config_t> inode_config;
std::map<std::string, inode_t> inode_by_name;
json11::Json node_placement;
@@ -160,24 +152,22 @@ public:
json11::Json::object serialize_inode_cfg(inode_config_t *cfg);
etcd_kv_t parse_etcd_kv(const json11::Json & kv_json);
std::vector<std::string> get_addresses();
void etcd_call_oneshot(std::string etcd_address, std::string api, json11::Json payload, int timeout, std::function<void(std::string, json11::Json)> callback);
void etcd_call(std::string api, json11::Json payload, int timeout, int retries, int interval, std::function<void(std::string, json11::Json)> callback);
virtual void etcd_call_oneshot(std::string etcd_address, std::string api, json11::Json payload, int timeout, std::function<void(std::string, json11::Json)> callback) = 0;
virtual void etcd_call(std::string api, json11::Json payload, int timeout, int retries, int interval, std::function<void(std::string, json11::Json)> callback) = 0;
void etcd_txn(json11::Json txn, int timeout, int retries, int interval, std::function<void(std::string, json11::Json)> callback);
void etcd_txn_slow(json11::Json txn, std::function<void(std::string, json11::Json)> callback);
void start_etcd_watcher();
void stop_ws_keepalive();
void start_ws_keepalive();
void load_global_config();
void load_pgs();
void reset_pg_exists();
void clean_nonexistent_pgs();
virtual void etcd_add_watch(json11::Json watch) = 0;
void load_global_config(std::function<void(const std::string &)> cb);
virtual void load_global_config() = 0;
void load_pgs(std::function<void(const std::string &)> cb);
virtual void load_pgs() = 0;
void parse_state(const etcd_kv_t & kv);
void parse_config(const json11::Json & config);
virtual void parse_config(const json11::Json & config);
void insert_inode_config(const inode_config_t & cfg);
inode_watch_t* watch_inode(std::string name);
void close_watch(inode_watch_t* watch);
int address_count();
~etcd_state_client_t();
virtual ~etcd_state_client_t();
static uint32_t parse_immediate_commit(const std::string & immediate_commit_str, uint32_t default_value);
static uint32_t parse_scheme(const std::string & scheme_str);
+487
View File
@@ -0,0 +1,487 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 or GNU GPL-2.0+ (see README.md for details)
#include "etcd_state_client_http.h"
#include "addr_util.h"
#include "http_client.h"
#include "str_util.h"
etcd_state_client_http_t::etcd_state_client_http_t(timerfd_manager_t *tfd)
{
this->tfd = tfd;
}
etcd_state_client_http_t::~etcd_state_client_http_t()
{
stop_ws_keepalive();
if (etcd_watch_ws)
{
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
}
if (keepalive_client)
{
http_close(keepalive_client);
keepalive_client = NULL;
}
if (load_pgs_timer_id >= 0)
{
tfd->clear_timer(load_pgs_timer_id);
load_pgs_timer_id = -1;
}
etcd_watches_initialised = -1;
}
void etcd_state_client_http_t::etcd_add_watch(json11::Json watch)
{
if (etcd_watch_ws)
{
http_post_message(etcd_watch_ws, WS_TEXT, watch.dump());
}
}
void etcd_state_client_http_t::etcd_call_oneshot(std::string etcd_address, std::string api, json11::Json payload,
int timeout, std::function<void(std::string, json11::Json)> callback)
{
std::string etcd_api_path;
int pos = etcd_address.find('/');
if (pos >= 0)
{
etcd_api_path = etcd_address.substr(pos);
etcd_address = etcd_address.substr(0, pos);
}
std::string req = payload.dump();
req = "POST "+etcd_api_path+api+" HTTP/1.1\r\n"
"Host: "+etcd_address+"\r\n"
"Content-Type: application/json\r\n"
"Content-Length: "+std::to_string(req.size())+"\r\n"
"Connection: close\r\n"
"\r\n"+req;
auto http_cli = http_init(tfd);
auto cb = [http_cli, callback](const http_response_t *response)
{
std::string err;
json11::Json data;
response->parse_json_response(err, data);
callback(err, data);
http_close(http_cli);
};
http_request(http_cli, etcd_address, req, { .timeout = timeout }, cb);
}
void etcd_state_client_http_t::etcd_call(std::string api, json11::Json payload, int timeout,
int retries, int interval, std::function<void(std::string, json11::Json)> callback)
{
if (!etcd_addresses.size() && !etcd_local.size())
{
fprintf(stderr, "etcd_address is missing in Vitastor configuration\n");
exit(1);
}
pick_next_etcd();
std::string etcd_address = selected_etcd_address;
std::string etcd_api_path;
int pos = etcd_address.find('/');
if (pos >= 0)
{
etcd_api_path = etcd_address.substr(pos);
etcd_address = etcd_address.substr(0, pos);
}
std::string req = payload.dump();
req = "POST "+etcd_api_path+api+" HTTP/1.1\r\n"
"Host: "+etcd_address+"\r\n"
"Content-Type: application/json\r\n"
"Content-Length: "+std::to_string(req.size())+"\r\n"
"Connection: keep-alive\r\n"
"Keep-Alive: timeout="+std::to_string(etcd_keepalive_timeout)+"\r\n"
"\r\n"+req;
retries--;
auto cb = [this, api, payload, timeout, retries, interval, callback,
cur_addr = selected_etcd_address](const http_response_t *response)
{
std::string err;
json11::Json data;
response->parse_json_response(err, data);
if (err != "")
{
if (cur_addr == selected_etcd_address)
selected_etcd_address = "";
if (retries > 0)
{
if (this->log_level > 0)
{
fprintf(
stderr, "Warning: etcd request failed: %s, retrying %d more times\n",
err.c_str(), retries
);
}
if (interval > 0)
{
// FIXME: Prevent destruction of etcd_state_client if timers or requests are active
tfd->set_timer(interval, false, [this, api, payload, timeout, retries, interval, callback](int)
{
etcd_call(api, payload, timeout, retries, interval, callback);
});
}
else
etcd_call(api, payload, timeout, retries, interval, callback);
}
else
callback(err, data);
}
else
callback(err, data);
};
if (!keepalive_client)
{
keepalive_client = http_init(tfd);
}
http_request(keepalive_client, etcd_address, req, { .timeout = timeout, .keepalive = true }, cb);
}
void etcd_state_client_http_t::parse_config(const json11::Json & config)
{
auto old_etcd_ws_keepalive_interval = this->etcd_ws_keepalive_interval;
etcd_state_client_t::parse_config(config);
if (this->etcd_ws_keepalive_interval != old_etcd_ws_keepalive_interval && ws_keepalive_timer >= 0)
{
stop_ws_keepalive();
start_ws_keepalive();
}
}
void etcd_state_client_http_t::pick_next_etcd()
{
if (selected_etcd_address != "")
return;
if (addresses_to_try.size() == 0)
{
// Prefer local etcd, if any
for (int i = 0; i < etcd_local.size(); i++)
addresses_to_try.push_back(etcd_local[i]);
std::vector<int> ns;
for (int i = 0; i < etcd_addresses.size(); i++)
ns.push_back(i);
if (!rand_initialized)
{
timespec tv;
clock_gettime(CLOCK_REALTIME, &tv);
srand48(tv.tv_sec*1000000000 + tv.tv_nsec);
rand_initialized = true;
}
while (ns.size())
{
int i = lrand48() % ns.size();
addresses_to_try.push_back(etcd_addresses[ns[i]]);
ns.erase(ns.begin()+i, ns.begin()+i+1);
}
}
selected_etcd_address = addresses_to_try[0];
addresses_to_try.erase(addresses_to_try.begin(), addresses_to_try.begin()+1);
}
void etcd_state_client_http_t::start_etcd_watcher()
{
if (!etcd_addresses.size() && !etcd_local.size())
{
fprintf(stderr, "etcd_address is missing in Vitastor configuration\n");
exit(1);
}
pick_next_etcd();
std::string etcd_address = selected_etcd_address;
std::string etcd_api_path;
int pos = etcd_address.find('/');
if (pos >= 0)
{
etcd_api_path = etcd_address.substr(pos);
etcd_address = etcd_address.substr(0, pos);
}
etcd_watches_initialised = 0;
ws_alive = 1;
if (etcd_watch_ws)
{
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
}
if (this->log_level > 1)
{
fprintf(stderr, "Trying to connect to etcd websocket at %s, watch from revision %ju/%ju/%ju\n", etcd_address.c_str(),
etcd_watch_revision_config, etcd_watch_revision_osd, etcd_watch_revision_pg);
}
etcd_watch_ws = open_websocket(tfd, etcd_address, etcd_api_path+"/watch", etcd_slow_timeout,
[this, cur_addr = selected_etcd_address](const http_response_t *msg)
{
if (msg->body.length())
{
ws_alive = 1;
std::string json_err;
json11::Json data = json11::Json::parse(msg->body, json_err);
if (json_err != "")
{
fprintf(stderr, "Bad JSON in etcd event: %s, ignoring event\n", json_err.c_str());
}
else
{
uint64_t watch_id = data["result"]["watch_id"].uint64_value();
if (data["result"]["created"].bool_value())
{
if (watch_id == ETCD_CONFIG_WATCH_ID ||
watch_id == ETCD_PG_STATE_WATCH_ID ||
watch_id == ETCD_OSD_STATE_WATCH_ID)
{
etcd_watches_initialised++;
}
if (etcd_watches_initialised == ETCD_TOTAL_WATCHES && this->log_level > 0)
{
fprintf(stderr, "Successfully subscribed to etcd at %s, revision %ju/%ju/%ju\n", cur_addr.c_str(),
etcd_watch_revision_config, etcd_watch_revision_osd, etcd_watch_revision_pg);
}
}
if (data["result"]["canceled"].bool_value())
{
// etcd watch canceled, maybe because the revision was compacted
if (data["result"]["compact_revision"].uint64_value())
{
// we may miss events if we proceed
// so we should restart from the beginning if we can
if (on_reload_hook != NULL)
{
// check to not trigger on_reload_hook multiple times
if (etcd_watch_ws != NULL)
{
fprintf(stderr, "Revisions before %ju were compacted by etcd, reloading state\n",
data["result"]["compact_revision"].uint64_value());
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
etcd_watch_revision_config = etcd_watch_revision_osd = etcd_watch_revision_pg = 0;
on_reload_hook();
}
return;
}
else
{
fprintf(stderr, "Revisions before %ju were compacted by etcd, exiting\n",
data["result"]["compact_revision"].uint64_value());
exit(1);
}
}
else
{
fprintf(stderr, "Watch canceled by etcd, reason: %s, exiting\n", data["result"]["cancel_reason"].string_value().c_str());
exit(1);
}
}
// Save revision only if it's present in the message - because sometimes etcd sends something without a header, like:
// {"error": {"grpc_code": 14, "http_code": 503, "http_status": "Service Unavailable", "message": "error reading from server: EOF"}}
// Also don't save revision from the initial created: true messages because they always contain the latest revision
if (etcd_watches_initialised == ETCD_TOTAL_WATCHES &&
!data["result"]["header"]["revision"].is_null() &&
!data["result"]["created"].bool_value())
{
// Restart watchers from the same revision number as in the last received message,
// not from the next one to protect against revision being split into multiple messages,
// even though etcd guarantees not to do that **within a single watcher** without fragment=true:
// https://etcd.io/docs/v3.5/learning/api_guarantees/#watch-apis
// Revision contents are ALWAYS split into separate messages for different watchers though!
// So generally we have to resume each watcher from its own revision...
// Progress messages may have watch_id=-1 if sent on behalf of multiple watchers though.
// And antietcd has an advanced semantic which merges the same revision for all watchers
// into one message and just omits watch_id.
// So we also have to handle the case where watch_id is -1 or not present (0).
auto watch_rev = data["result"]["header"]["revision"].uint64_value();
if (!watch_id || watch_id == UINT64_MAX)
etcd_watch_revision_config = etcd_watch_revision_osd = etcd_watch_revision_pg = watch_rev;
else if (watch_id == ETCD_CONFIG_WATCH_ID)
etcd_watch_revision_config = watch_rev;
else if (watch_id == ETCD_PG_STATE_WATCH_ID)
etcd_watch_revision_pg = watch_rev;
else if (watch_id == ETCD_OSD_STATE_WATCH_ID)
etcd_watch_revision_osd = watch_rev;
addresses_to_try.clear();
}
// First gather all changes into a hash to remove multiple overwrites
std::map<std::string, etcd_kv_t> changes;
for (auto & ev: data["result"]["events"].array_items())
{
auto kv = parse_etcd_kv(ev["kv"]);
if (kv.key != "")
{
changes[kv.key] = kv;
}
}
for (auto & kv: changes)
{
if (this->log_level > 3)
{
fprintf(stderr, "Incoming event: %s -> %s\n", kv.first.c_str(), kv.second.value.dump().c_str());
}
parse_state(kv.second);
}
// React to changes
if (on_change_hook != NULL)
{
on_change_hook(changes);
}
}
}
if (msg->eof)
{
fprintf(stderr, "Disconnected from etcd %s\n", cur_addr.c_str());
if (cur_addr == selected_etcd_address)
selected_etcd_address = "";
if (etcd_watch_ws)
{
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
}
if (etcd_watches_initialised == 0)
{
// Connection not established, retry in <etcd_quick_timeout>
tfd->set_timer(etcd_quick_timeout, false, [this](int)
{
start_etcd_watcher();
});
}
else if (etcd_watches_initialised > 0)
{
// Connection was live, retry immediately
etcd_watches_initialised = 0;
start_etcd_watcher();
}
}
});
http_post_message(etcd_watch_ws, WS_TEXT, json11::Json(json11::Json::object {
{ "create_request", json11::Json::object {
{ "key", base64_encode(etcd_prefix+"/config/") },
{ "range_end", base64_encode(etcd_prefix+"/config0") },
{ "start_revision", etcd_watch_revision_config },
{ "watch_id", ETCD_CONFIG_WATCH_ID },
{ "progress_notify", true },
} }
}).dump());
http_post_message(etcd_watch_ws, WS_TEXT, json11::Json(json11::Json::object {
{ "create_request", json11::Json::object {
{ "key", base64_encode(etcd_prefix+"/osd/state/") },
{ "range_end", base64_encode(etcd_prefix+"/osd/state0") },
{ "start_revision", etcd_watch_revision_osd },
{ "watch_id", ETCD_OSD_STATE_WATCH_ID },
{ "progress_notify", true },
} }
}).dump());
http_post_message(etcd_watch_ws, WS_TEXT, json11::Json(json11::Json::object {
{ "create_request", json11::Json::object {
{ "key", base64_encode(etcd_prefix+"/pg/") },
{ "range_end", base64_encode(etcd_prefix+"/pg0") },
{ "start_revision", etcd_watch_revision_pg },
{ "watch_id", ETCD_PG_STATE_WATCH_ID },
{ "progress_notify", true },
} }
}).dump());
// FIXME: Do not watch /pg/history/ at all in client code (not in OSD)
if (on_start_watcher_hook)
{
on_start_watcher_hook(etcd_watch_ws);
}
start_ws_keepalive();
}
void etcd_state_client_http_t::stop_ws_keepalive()
{
if (ws_keepalive_timer >= 0)
{
tfd->clear_timer(ws_keepalive_timer);
ws_keepalive_timer = -1;
}
}
void etcd_state_client_http_t::start_ws_keepalive()
{
if (ws_keepalive_timer < 0)
{
ws_keepalive_timer = tfd->set_timer(etcd_ws_keepalive_interval*1000, true, [this](int)
{
if (!etcd_watch_ws || etcd_watches_initialised < ETCD_TOTAL_WATCHES)
{
// Do nothing
}
else if (!ws_alive)
{
if (this->log_level > 0)
{
fprintf(stderr, "Websocket ping failed, disconnecting from etcd %s\n", selected_etcd_address.c_str());
}
if (etcd_watch_ws)
{
http_close(etcd_watch_ws);
etcd_watch_ws = NULL;
}
start_etcd_watcher();
}
else
{
ws_alive = 0;
http_post_message(etcd_watch_ws, WS_TEXT, json11::Json(json11::Json::object {
{ "progress_request", json11::Json::object { } }
}).dump());
}
});
}
}
void etcd_state_client_http_t::load_global_config()
{
etcd_state_client_t::load_global_config([this](const std::string & err)
{
if (err != "")
{
fprintf(stderr, "Error reading configuration from etcd: %s\n", err.c_str());
if (infinite_start)
{
tfd->set_timer(etcd_slow_timeout, false, [this](int timer_id)
{
load_global_config();
});
}
else
{
exit(1);
}
}
});
}
void etcd_state_client_http_t::load_pgs()
{
timespec tv;
clock_gettime(CLOCK_REALTIME, &tv);
uint64_t ms_passed = (tv.tv_sec-etcd_last_reload.tv_sec)*1000 + (tv.tv_nsec-etcd_last_reload.tv_nsec)/1000000;
if (ms_passed < etcd_min_reload_interval)
{
if (load_pgs_timer_id < 0)
{
load_pgs_timer_id = tfd->set_timer(etcd_min_reload_interval+50-ms_passed, false, [this](int) { load_pgs(); });
}
return;
}
etcd_last_reload = tv;
if (load_pgs_timer_id >= 0)
{
tfd->clear_timer(load_pgs_timer_id);
load_pgs_timer_id = -1;
}
etcd_state_client_t::load_pgs([this](const std::string & err)
{
if (err != "")
{
// Retry indefinitely
fprintf(stderr, "Error loading PGs from etcd: %s\n", err.c_str());
tfd->set_timer(etcd_slow_timeout, false, [this](int timer_id)
{
load_pgs();
});
}
else
{
start_etcd_watcher();
}
});
}
+36
View File
@@ -0,0 +1,36 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 or GNU GPL-2.0+ (see README.md for details)
#pragma once
#include "etcd_state_client.h"
struct __attribute__((visibility("default"))) etcd_state_client_http_t: public etcd_state_client_t
{
protected:
timerfd_manager_t *tfd = NULL;
std::string selected_etcd_address;
std::vector<std::string> addresses_to_try;
int ws_keepalive_timer = -1;
int ws_alive = 0;
bool rand_initialized = false;
int etcd_watches_initialised = 0;
timespec etcd_last_reload = {};
int load_pgs_timer_id = -1;
http_co_t *keepalive_client = NULL;
void pick_next_etcd();
void start_etcd_watcher();
void stop_ws_keepalive();
void start_ws_keepalive();
public:
http_co_t *etcd_watch_ws = NULL;
etcd_state_client_http_t(timerfd_manager_t *tfd);
void etcd_call_oneshot(std::string etcd_address, std::string api, json11::Json payload, int timeout, std::function<void(std::string, json11::Json)> callback) override;
void etcd_call(std::string api, json11::Json payload, int timeout, int retries, int interval, std::function<void(std::string, json11::Json)> callback) override;
void etcd_add_watch(json11::Json watch) override;
void load_global_config() override;
void load_pgs() override;
void parse_config(const json11::Json & config) override;
~etcd_state_client_http_t();
};
+205
View File
@@ -0,0 +1,205 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 or GNU GPL-2.0+ (see README.md for details)
#include <assert.h>
#include "etcd_state_client_mock.h"
#include "str_util.h"
etcd_state_client_mock_t::etcd_state_client_mock_t()
{
timespec tv;
clock_gettime(CLOCK_REALTIME, &tv);
srand48(tv.tv_sec*1000000000 + tv.tv_nsec);
}
void etcd_state_client_mock_t::etcd_add_watch(json11::Json watch)
{
}
void etcd_state_client_mock_t::etcd_call_oneshot(std::string etcd_address, std::string api, json11::Json payload,
int timeout, std::function<void(std::string, json11::Json)> callback)
{
}
void etcd_state_client_mock_t::pause()
{
paused = true;
}
void etcd_state_client_mock_t::resume()
{
paused = false;
auto queue = std::move(this->queue);
for (auto& req: queue)
{
etcd_call(req.api, req.payload, req.timeout, req.retries, req.interval, req.callback);
}
}
void etcd_state_client_mock_t::set(const std::string& key, json11::Json data, uint64_t mod_revision, uint64_t lease_id)
{
if (!mod_revision)
mod_revision = ++this->mod_revision;
this->data[key] = (etcd_mock_key_data_t){ .value = data.dump(), .mod_revision = mod_revision, .lease_id = lease_id };
}
void etcd_state_client_mock_t::etcd_call(std::string api, json11::Json payload, int timeout,
int retries, int interval, std::function<void(std::string, json11::Json)> callback)
{
if (paused)
{
queue.push_back({ api, payload, timeout, retries, interval, callback });
return;
}
printf("+ etcd: %s\n", api.c_str());
if (api == "/kv/txn")
{
bool ok = true;
for (auto& check: payload["compare"].array_items())
{
auto key = base64_decode(check["key"].string_value());
etcd_mock_key_data_t *key_data = data.find(key) != data.end() ? &data.at(key) : NULL;
auto target = check["target"].string_value();
auto res = check["result"].string_value();
assert(res == "LESS" || res == "");
bool less = res == "LESS";
if (target == "MOD")
{
uint64_t rev = check["mod_revision"].uint64_value();
assert(!less || rev);
ok = ok && (less ? (!key_data || key_data->mod_revision < rev) : (key_data && key_data->mod_revision == rev));
}
else if (target == "CREATE")
{
uint64_t rev = check["create_revision"].uint64_value();
assert(rev == 0 && !less);
ok = ok && !key_data;
}
else if (target == "VERSION")
{
uint64_t rev = check["version"].uint64_value();
assert(rev == 0 && !less);
ok = ok && !key_data;
}
else if (target == "LEASE")
{
assert(!less);
uint64_t lease_id = check["lease"].uint64_value();
ok = ok && key_data && key_data->lease_id == lease_id;
}
else
assert(0);
}
std::map<std::string, etcd_kv_t> changes;
bool has_mod = false;
for (auto& op: payload[ok ? "success" : "failure"].array_items())
{
auto& obj = op.object_items();
has_mod = has_mod || obj.find("request_put") != obj.end() ||
obj.find("request_delete_range") != obj.end();
}
if (has_mod)
{
mod_revision++;
}
json11::Json::array responses;
for (auto& op_ptr: payload[ok ? "success" : "failure"].array_items())
{
auto& op = op_ptr.object_items();
if (op.find("request_range") != op.end())
{
json11::Json::array kvs;
auto req = op.at("request_range");
auto key = base64_decode(req["key"].string_value());
auto range_end = base64_decode(req["range_end"].string_value());
auto begin_it = range_end.empty() ? data.find(key) : data.lower_bound(key);
auto end_it = range_end.empty() ? (begin_it == data.end() ? begin_it : std::next(begin_it)) : data.lower_bound(range_end);
for (auto it = begin_it; it != end_it; it++)
{
printf("\\- get: %s = %s, rev %ju\n", it->first.c_str(), it->second.value.c_str(), it->second.mod_revision);
kvs.push_back(json11::Json::object {
{ "key", base64_encode(it->first) },
{ "value", base64_encode(it->second.value) },
{ "mod_revision", it->second.mod_revision },
});
}
responses.push_back(json11::Json::object {
{ "response_range", json11::Json::object{ { "header", json11::Json::object{ { "revision", mod_revision } } }, { "kvs", kvs } } },
});
}
else if (op.find("request_put") != op.end())
{
auto req = op.at("request_put");
auto key = base64_decode(req["key"].string_value());
auto value = base64_decode(req["value"].string_value());
auto lease_id = req["lease"].uint64_value();
printf("\\- put: %s = %s, rev %ju, lease %ju\n", key.c_str(), value.c_str(), mod_revision, lease_id);
data[key] = {
.value = value,
.mod_revision = mod_revision,
.lease_id = lease_id,
};
std::string err;
json11::Json json_value = json11::Json::parse(value, err);
if (err != "")
{
fprintf(stderr, "Invalid JSON in etcd key %s during test: %s\n", key.c_str(), value.c_str());
exit(1);
}
changes[key] = { .key = key, .value = json_value, .mod_revision = mod_revision };
responses.push_back(json11::Json::object {
{ "response_put", json11::Json::object{ { "header", json11::Json::object{ { "revision", mod_revision } } } } },
});
}
else if (op.find("request_delete_range") != op.end())
{
auto req = op.at("request_delete_range");
auto key = base64_decode(req["key"].string_value());
auto range_end = base64_decode(req["range_end"].string_value());
uint64_t n_del = 0;
for (auto it = data.lower_bound(key); it != data.end() && (range_end == "" || it->first < range_end); )
{
auto & key = it->first;
printf("\\- del: %s\n", key.c_str());
changes[key] = { .key = key, .mod_revision = mod_revision };
n_del++;
data.erase(it++);
}
responses.push_back(json11::Json::object {
{ "response_delete_range", json11::Json::object{ { "header", json11::Json::object{ { "revision", mod_revision } } }, { "deleted", n_del } } },
});
}
}
callback("", json11::Json::object{
{ "header", json11::Json::object{ { "revision", mod_revision } } },
{ "succeeded", ok },
{ "responses", responses }
});
// Push changes to watcher
if (changes.size())
{
for (auto & kv: changes)
parse_state(kv.second);
if (on_change_hook != NULL)
on_change_hook(changes);
}
}
else if (api == "/lease/grant")
{
uint64_t lease_id = (((uint64_t)lrand48()) << 32) | lrand48();
leases[lease_id] = payload["TTL"].uint64_value();
callback("", json11::Json::object{ { "ID", std::to_string(lease_id) } });
}
else
callback("Unsupported", json11::Json());
}
void etcd_state_client_mock_t::load_global_config()
{
etcd_state_client_t::load_global_config([this](const std::string & err) {});
}
void etcd_state_client_mock_t::load_pgs()
{
etcd_state_client_t::load_pgs([this](const std::string & err) {});
}
+42
View File
@@ -0,0 +1,42 @@
// Copyright (c) Vitaliy Filippov, 2019+
// License: VNPL-1.1 or GNU GPL-2.0+ (see README.md for details)
#pragma once
#include "etcd_state_client.h"
struct etcd_mock_key_data_t
{
std::string value;
uint64_t mod_revision;
uint64_t lease_id;
};
struct etcd_mock_request_t
{
std::string api;
json11::Json payload;
int timeout;
int retries;
int interval;
std::function<void(std::string, json11::Json)> callback;
};
struct etcd_state_client_mock_t: public etcd_state_client_t
{
uint64_t mod_revision = 0;
bool paused = false;
std::vector<etcd_mock_request_t> queue;
public:
std::map<uint64_t, uint64_t> leases;
std::map<std::string, etcd_mock_key_data_t> data;
etcd_state_client_mock_t();
void set(const std::string& key, json11::Json data, uint64_t mod_revision = 0, uint64_t lease_id = 0);
void pause();
void resume();
void etcd_call_oneshot(std::string etcd_address, std::string api, json11::Json payload, int timeout, std::function<void(std::string, json11::Json)> callback) override;
void etcd_call(std::string api, json11::Json payload, int timeout, int retries, int interval, std::function<void(std::string, json11::Json)> callback) override;
void etcd_add_watch(json11::Json watch) override;
void load_global_config() override;
void load_pgs() override;
};
+9 -4
View File
@@ -38,6 +38,8 @@
#define MSGR_SENDP_HDR 1
#define MSGR_SENDP_FREE 2
#define MAX_SIMPLE_PAYLOAD_SIZE 1048576
struct msgr_sendp_t
{
osd_op_t *op;
@@ -181,9 +183,12 @@ public:
timerfd_manager_t *tfd = NULL;
ring_loop_i *ringloop = NULL;
bool has_sendmsg_zc = false;
// osd_num_t is only for logging and asserts
uint64_t next_client_id = 1;
osd_num_t osd_num;
// osd_num = 0 for client messenger, osd_num > 0 for OSD messenger
osd_num_t osd_num = 0;
uint32_t clean_entry_bitmap_size = 0;
uint32_t bs_block_size = 0;
uint32_t max_write_request_size = 0;
robin_hood::unordered_flat_map<uint64_t, osd_client_t*> clients;
robin_hood::unordered_flat_map<uint64_t, osd_client_t*> osd_peers;
robin_hood::unordered_flat_map<int, osd_client_t*> clients_by_fd;
@@ -222,7 +227,7 @@ public:
#ifdef WITH_RDMA
bool is_rdma_enabled();
bool connect_rdma(uint64_t client_id, std::string rdma_address, uint64_t client_max_msg);
json11::Json connect_rdma(uint64_t client_id, std::string rdma_address, uint64_t client_max_msg);
#endif
#ifdef WITH_RDMACM
bool is_use_rdmacm();
@@ -249,7 +254,7 @@ protected:
bool handle_read(int result, osd_client_t *cl);
bool handle_read_buffer(osd_client_t *cl, void *curbuf, int remain);
bool handle_finished_read(osd_client_t *cl);
void handle_op_hdr(osd_client_t *cl);
bool handle_op_hdr(osd_client_t *cl);
bool handle_reply_hdr(osd_client_t *cl);
void handle_reply_ready(osd_op_t *op);
void handle_immediate_ops();
+50
View File
@@ -3,6 +3,7 @@
#include <assert.h>
#include "messenger.h"
#include "msgr_op.h"
osd_op_t::~osd_op_t()
@@ -38,3 +39,52 @@ bool osd_op_t::is_recovery_related()
req.hdr.opcode == OSD_OP_SEC_SYNC &&
(req.sec_sync.flags & OSD_OP_RECOVERY_RELATED);
}
void osd_messenger_t::measure_exec(osd_op_t *cur_op)
{
// Measure execution latency
if (cur_op->req.hdr.opcode > OSD_OP_MAX)
{
return;
}
if (!cur_op->tv_end.tv_sec)
{
clock_gettime(CLOCK_REALTIME, &cur_op->tv_end);
}
uint64_t len = 0;
if (cur_op->req.hdr.opcode == OSD_OP_READ ||
cur_op->req.hdr.opcode == OSD_OP_WRITE ||
cur_op->req.hdr.opcode == OSD_OP_SCRUB)
{
// req.rw.len is internally set to the full object size for scrubs
len = cur_op->req.rw.len;
}
else if (cur_op->req.hdr.opcode == OSD_OP_SEC_READ ||
cur_op->req.hdr.opcode == OSD_OP_SEC_WRITE ||
cur_op->req.hdr.opcode == OSD_OP_SEC_WRITE_STABLE)
{
len = cur_op->req.sec_rw.len;
}
inc_op_stats(stats, cur_op->req.hdr.opcode, cur_op->tv_begin, cur_op->tv_end, len);
if (cur_op->is_recovery_related())
{
inc_op_stats(recovery_stats, cur_op->req.hdr.opcode, cur_op->tv_begin, cur_op->tv_end, len);
}
}
void osd_messenger_t::inc_op_stats(osd_op_stats_t & stats, uint64_t opcode, timespec & tv_begin, timespec & tv_end, uint64_t len)
{
uint64_t usecs = (
(tv_end.tv_sec - tv_begin.tv_sec)*1000000 +
(tv_end.tv_nsec - tv_begin.tv_nsec)/1000
);
stats.op_stat_count[opcode]++;
if (!stats.op_stat_count[opcode])
{
stats.op_stat_count[opcode] = 1;
stats.op_stat_sum[opcode] = 0;
stats.op_stat_bytes[opcode] = 0;
}
stats.op_stat_sum[opcode] += usecs;
stats.op_stat_bytes[opcode] += len;
}
+1 -1
View File
@@ -165,7 +165,7 @@ struct __attribute__((visibility("default"))) osd_op_t
// bitmap, bitmap_len, bmp_data are only meaningful for reads
void *bitmap = NULL;
unsigned bitmap_len = 0;
unsigned bmp_data = 0;
size_t bmp_data = 0;
void *bitmap_buf = NULL;
void *rmw_buf = NULL;
osd_primary_op_data_t* op_data = NULL;
+7 -4
View File
@@ -507,7 +507,7 @@ int msgr_rdma_connection_t::connect(msgr_rdma_address_t *dest)
return 0;
}
bool osd_messenger_t::connect_rdma(uint64_t client_id, std::string rdma_address, uint64_t client_max_msg)
json11::Json osd_messenger_t::connect_rdma(uint64_t client_id, std::string rdma_address, uint64_t client_max_msg)
{
// Try to connect to the peer using RDMA
msgr_rdma_address_t addr;
@@ -523,7 +523,7 @@ bool osd_messenger_t::connect_rdma(uint64_t client_id, std::string rdma_address,
{
if (log_level > 0)
fprintf(stderr, "No RDMA context for peer %ju, using only TCP\n", client_id);
return false;
return json11::Json();
}
msgr_rdma_connection_t *rdma_conn = msgr_rdma_connection_t::create(selected_ctx, rdma_max_send, rdma_max_recv, rdma_max_sge, client_max_msg);
if (rdma_conn)
@@ -542,11 +542,14 @@ bool osd_messenger_t::connect_rdma(uint64_t client_id, std::string rdma_address,
// Remember connection, but switch to RDMA only after sending the configuration response
cl->rdma_conn = rdma_conn;
cl->peer_state = PEER_RDMA_CONNECTING;
return true;
return json11::Json::object{
{"rdma_address", rdma_conn->addr.to_string()},
{"rdma_max_msg", rdma_conn->max_msg},
};
}
}
}
return false;
return json11::Json();
}
static void try_send_rdma_wr(osd_client_t *cl, ibv_sge *sge, int op_sge)
+62 -3
View File
@@ -231,7 +231,11 @@ bool osd_messenger_t::handle_finished_read(osd_client_t *cl)
}
cl->read_op_id++;
}
handle_op_hdr(cl);
if (!handle_op_hdr(cl))
{
stop_client(cl->client_id);
return false;
}
}
else
{
@@ -262,7 +266,7 @@ bool osd_messenger_t::handle_finished_read(osd_client_t *cl)
return true;
}
void osd_messenger_t::handle_op_hdr(osd_client_t *cl)
bool osd_messenger_t::handle_op_hdr(osd_client_t *cl)
{
osd_op_t *cur_op = cl->read_op;
if (cur_op->req.hdr.opcode == OSD_OP_SEC_READ)
@@ -274,7 +278,16 @@ void osd_messenger_t::handle_op_hdr(osd_client_t *cl)
{
if (cur_op->req.sec_rw.attr_len > 0)
{
if (cur_op->req.sec_rw.attr_len > sizeof(unsigned))
if (cur_op->req.sec_rw.attr_len > clean_entry_bitmap_size)
{
if (log_level > 1)
{
fprintf(stderr, "Error: peer %ju secondary write request attr_len too large (%u > %u bytes), stopping\n", cl->client_id,
cur_op->req.sec_rw.attr_len, clean_entry_bitmap_size);
}
return false;
}
else if (cur_op->req.sec_rw.attr_len > sizeof(cur_op->bmp_data))
cur_op->bitmap = cur_op->rmw_buf = malloc_or_die(cur_op->req.sec_rw.attr_len);
else
cur_op->bitmap = &cur_op->bmp_data;
@@ -282,6 +295,15 @@ void osd_messenger_t::handle_op_hdr(osd_client_t *cl)
}
if (cur_op->req.sec_rw.len > 0)
{
if (cur_op->req.sec_rw.len > bs_block_size)
{
if (log_level > 1)
{
fprintf(stderr, "Error: peer %ju secondary write request size too large (%u > %u bytes), stopping\n", cl->client_id,
cur_op->req.sec_rw.len, bs_block_size);
}
return false;
}
cur_op->buf = memalign_or_die(MEM_ALIGNMENT, cur_op->req.sec_rw.len);
cl->recv_list.push_back(cur_op->buf, cur_op->req.sec_rw.len);
}
@@ -292,6 +314,15 @@ void osd_messenger_t::handle_op_hdr(osd_client_t *cl)
{
if (cur_op->req.sec_stab.len > 0)
{
if (cur_op->req.sec_stab.len > MAX_SIMPLE_PAYLOAD_SIZE)
{
if (log_level > 1)
{
fprintf(stderr, "Error: peer %ju stabilize request size too large (%lu > %u bytes), stopping\n", cl->client_id,
cur_op->req.sec_stab.len, MAX_SIMPLE_PAYLOAD_SIZE);
}
return false;
}
cur_op->buf = memalign_or_die(MEM_ALIGNMENT, cur_op->req.sec_stab.len);
cl->recv_list.push_back(cur_op->buf, cur_op->req.sec_stab.len);
}
@@ -301,6 +332,15 @@ void osd_messenger_t::handle_op_hdr(osd_client_t *cl)
{
if (cur_op->req.sec_read_bmp.len > 0)
{
if (cur_op->req.sec_read_bmp.len > MAX_SIMPLE_PAYLOAD_SIZE)
{
if (log_level > 1)
{
fprintf(stderr, "Error: peer %ju sec_read_bmp request size too large (%lu > %u bytes), stopping\n", cl->client_id,
cur_op->req.sec_read_bmp.len, MAX_SIMPLE_PAYLOAD_SIZE);
}
return false;
}
cur_op->buf = memalign_or_die(MEM_ALIGNMENT, cur_op->req.sec_read_bmp.len);
cl->recv_list.push_back(cur_op->buf, cur_op->req.sec_read_bmp.len);
}
@@ -310,6 +350,15 @@ void osd_messenger_t::handle_op_hdr(osd_client_t *cl)
{
if (cur_op->req.rw.len > 0)
{
if (cur_op->req.rw.len > max_write_request_size)
{
if (log_level > 1)
{
fprintf(stderr, "Error: peer %ju write request size too large (%u > %u bytes), stopping\n", cl->client_id,
cur_op->req.rw.len, max_write_request_size);
}
return false;
}
cur_op->buf = memalign_or_die(MEM_ALIGNMENT, cur_op->req.rw.len);
cl->recv_list.push_back(cur_op->buf, cur_op->req.rw.len);
}
@@ -319,6 +368,15 @@ void osd_messenger_t::handle_op_hdr(osd_client_t *cl)
{
if (cur_op->req.show_conf.json_len > 0)
{
if (cur_op->req.show_conf.json_len > MAX_SIMPLE_PAYLOAD_SIZE)
{
if (log_level > 1)
{
fprintf(stderr, "Error: peer %ju show_config request length too large (%lu > %u bytes), stopping\n", cl->client_id,
cur_op->req.show_conf.json_len, MAX_SIMPLE_PAYLOAD_SIZE);
}
return false;
}
cur_op->buf = malloc_or_die(cur_op->req.show_conf.json_len+1);
((uint8_t*)cur_op->buf)[cur_op->req.show_conf.json_len] = 0;
cl->recv_list.push_back(cur_op->buf, cur_op->req.show_conf.json_len);
@@ -344,6 +402,7 @@ void osd_messenger_t::handle_op_hdr(osd_client_t *cl)
cl->read_op = NULL;
cl->read_state = 0;
}
return true;
}
bool osd_messenger_t::handle_reply_hdr(osd_client_t *cl)
-49
View File
@@ -140,55 +140,6 @@ void osd_messenger_t::outbox_push(osd_op_t *cur_op)
}
}
void osd_messenger_t::inc_op_stats(osd_op_stats_t & stats, uint64_t opcode, timespec & tv_begin, timespec & tv_end, uint64_t len)
{
uint64_t usecs = (
(tv_end.tv_sec - tv_begin.tv_sec)*1000000 +
(tv_end.tv_nsec - tv_begin.tv_nsec)/1000
);
stats.op_stat_count[opcode]++;
if (!stats.op_stat_count[opcode])
{
stats.op_stat_count[opcode] = 1;
stats.op_stat_sum[opcode] = 0;
stats.op_stat_bytes[opcode] = 0;
}
stats.op_stat_sum[opcode] += usecs;
stats.op_stat_bytes[opcode] += len;
}
void osd_messenger_t::measure_exec(osd_op_t *cur_op)
{
// Measure execution latency
if (cur_op->req.hdr.opcode > OSD_OP_MAX)
{
return;
}
if (!cur_op->tv_end.tv_sec)
{
clock_gettime(CLOCK_REALTIME, &cur_op->tv_end);
}
uint64_t len = 0;
if (cur_op->req.hdr.opcode == OSD_OP_READ ||
cur_op->req.hdr.opcode == OSD_OP_WRITE ||
cur_op->req.hdr.opcode == OSD_OP_SCRUB)
{
// req.rw.len is internally set to the full object size for scrubs
len = cur_op->req.rw.len;
}
else if (cur_op->req.hdr.opcode == OSD_OP_SEC_READ ||
cur_op->req.hdr.opcode == OSD_OP_SEC_WRITE ||
cur_op->req.hdr.opcode == OSD_OP_SEC_WRITE_STABLE)
{
len = cur_op->req.sec_rw.len;
}
inc_op_stats(stats, cur_op->req.hdr.opcode, cur_op->tv_begin, cur_op->tv_end, len);
if (cur_op->is_recovery_related())
{
inc_op_stats(recovery_stats, cur_op->req.hdr.opcode, cur_op->tv_begin, cur_op->tv_end, len);
}
}
bool osd_messenger_t::try_send(osd_client_t *cl)
{
if (!cl->send_list.size() || cl->write_msg.msg_iovlen > 0 || cl->peer_state == PEER_STOPPED || cl->peer_fd < 0)
+12 -11
View File
@@ -301,10 +301,9 @@ const char *help_text =
" --nbd_disconnect_on_close 1\n"
" Disconnect the nbd device on close by last opener.\n"
#endif
#ifdef NBD_FLAG_READ_ONLY
" --readonly\n"
" --nbd_ro 1\n"
" Set device into read only mode.\n"
#endif
"\n"
"vitastor-nbd netlink-unmap /dev/nbdN\n"
" Unmap a device using netlink interface. Works with both netlink and ioctl mapped devices.\n"
@@ -348,6 +347,7 @@ protected:
int read_ready = 0;
msghdr read_msg = { 0 }, send_msg = { 0 };
iovec read_iov = { 0 };
bool stop = false;
std::string logfile = "/dev/null";
@@ -514,7 +514,7 @@ help:
// Create client
ringloop = new ring_loop_t(RINGLOOP_DEFAULT_SIZE);
epmgr = new epoll_manager_t(ringloop);
cli = new cluster_client_t(ringloop, epmgr->tfd, cfg);
cli = cluster_client_t::create(ringloop, epmgr->tfd, cfg);
if (!inode)
{
// Load image metadata
@@ -525,7 +525,7 @@ help:
break;
ringloop->wait();
}
watch = cli->st_cli.watch_inode(image_name);
watch = cli->st_cli->watch_inode(image_name);
device_size = watch->cfg.size;
if (!watch->cfg.num || !device_size)
{
@@ -581,10 +581,8 @@ help:
}
uint64_t flags = NBD_FLAG_SEND_FLUSH;
uint64_t cflags = 0;
#ifdef NBD_FLAG_READ_ONLY
if (!cfg["nbd_ro"].is_null())
if (!cfg["readonly"].is_null() || !cfg["nbd_ro"].is_null())
flags |= NBD_FLAG_READ_ONLY;
#endif
#ifdef NBD_CFLAG_DESTROY_ON_DISCONNECT
if (!cfg["nbd_destroy_on_disconnect"].is_null())
cflags |= NBD_CFLAG_DESTROY_ON_DISCONNECT;
@@ -620,7 +618,10 @@ help:
if (!cfg["dev_num"].is_null())
{
int r;
if ((r = run_nbd(sockfd, cfg["dev_num"].int64_value(), device_size, NBD_FLAG_SEND_FLUSH, nbd_timeout, bg)) != 0)
uint64_t flags = NBD_FLAG_SEND_FLUSH;
if (!cfg["readonly"].is_null())
flags |= NBD_FLAG_READ_ONLY;
if ((r = run_nbd(sockfd, cfg["dev_num"].int64_value(), device_size, flags, nbd_timeout, bg)) != 0)
{
fprintf(stderr, "run_nbd: %s\n", strerror(-r));
exit(1);
@@ -678,8 +679,7 @@ help:
};
ringloop->register_consumer(&consumer);
// Add FD to epoll
bool stop = false;
epmgr->tfd->set_fd_handler(sockfd[0], false, [this, &stop](int peer_fd, int epoll_events)
epmgr->tfd->set_fd_handler(sockfd[0], false, [this](int peer_fd, int epoll_events)
{
if (epoll_events & EPOLLRDHUP)
{
@@ -1118,7 +1118,8 @@ protected:
{
// Disconnect
close(nbd_fd);
exit(0);
stop = true;
return;
}
if (be32toh(cur_req.magic) != NBD_REQUEST_MAGIC ||
req_type != NBD_CMD_READ && req_type != NBD_CMD_WRITE && req_type != NBD_CMD_FLUSH)
+76
View File
@@ -705,6 +705,54 @@ static void vitastor_close(BlockDriverState *bs)
client->last_bitmap = NULL;
}
// Unregister all event sources from the current AioContext. Called by the
// block layer before bs is moved to a different AioContext (e.g. during live
// migration, drain, dataplane switching). The block layer guarantees that no
// requests are in flight at this point.
static void vitastor_detach_aio_context(BlockDriverState *bs)
{
VitastorClient *client = bs->opaque;
int i;
#if defined VITASTOR_C_API_VERSION && VITASTOR_C_API_VERSION >= 2
if (client->uring_eventfd >= 0)
{
universal_aio_set_fd_handler(client->ctx, client->uring_eventfd, NULL, NULL, NULL);
// Wait until any scheduled B/H is processed before switching contexts:
// it would otherwise fire on the old context with stale state.
if (client->bh_uring_scheduled)
{
BDRV_POLL_WHILE(bs, client->bh_uring_scheduled);
}
}
#endif
for (i = 0; i < client->fd_count; i++)
{
universal_aio_set_fd_handler(client->ctx, client->fds[i]->fd, NULL, NULL, NULL);
}
}
// (Re-)register all event sources on the new AioContext.
static void vitastor_attach_aio_context(BlockDriverState *bs, AioContext *new_ctx)
{
VitastorClient *client = bs->opaque;
int i;
client->ctx = new_ctx;
#if defined VITASTOR_C_API_VERSION && VITASTOR_C_API_VERSION >= 2
if (client->uring_eventfd >= 0)
{
universal_aio_set_fd_handler(new_ctx, client->uring_eventfd, vitastor_uring_handler, NULL, client);
}
#endif
for (i = 0; i < client->fd_count; i++)
{
VitastorFdData *fdd = client->fds[i];
universal_aio_set_fd_handler(new_ctx, fdd->fd,
fdd->fd_read ? vitastor_aio_fd_read : NULL,
fdd->fd_write ? vitastor_aio_fd_write : NULL,
fdd);
}
}
#if QEMU_VERSION_MAJOR >= 3 || QEMU_VERSION_MAJOR == 2 && QEMU_VERSION_MINOR >= 2
static void vitastor_refresh_filename(BlockDriverState *bs)
{
@@ -832,6 +880,25 @@ static int vitastor_refresh_limits(BlockDriverState *bs)
// return 0;
//}
// Move the running coroutine to the BlockDriverState's home AioContext.
//
// The block-coroutine-wrapper generator sets poll_state.ctx to
// qemu_get_current_aio_context() in the sync wrappers (bdrv_flush(),
// bdrv_pread() etc.). When bdrv_flush_all() runs under BQL from outside the
// bs's iothread (e.g. on the migration thread inside do_vm_stop()), that is
// the main AioContext, not the iothread that actually owns the bs. The
// coroutine then runs on the wrong context while completions are delivered on
// the iothread, and racing aio_co_schedule() vs. qemu_aio_coroutine_enter()
// on the same coroutine triggers "Co-routine was already scheduled in
// aio_co_schedule" and aborts the process (observed during live migration).
//
// aio_co_reschedule_self() is a no-op when we are already on the target ctx.
#if QEMU_VERSION_MAJOR > 5 || QEMU_VERSION_MAJOR == 5 && QEMU_VERSION_MINOR >= 2
#define vitastor_co_pin_to_bs_ctx(bs) aio_co_reschedule_self(bdrv_get_aio_context(bs))
#else
#define vitastor_co_pin_to_bs_ctx(bs) ((void)0)
#endif
static void vitastor_co_init_task(BlockDriverState *bs, VitastorRPC *task)
{
*task = (VitastorRPC) {
@@ -890,6 +957,7 @@ static int coroutine_fn vitastor_co_preadv(BlockDriverState *bs,
{
VitastorClient *client = bs->opaque;
VitastorRPC task;
vitastor_co_pin_to_bs_ctx(bs);
vitastor_co_init_task(bs, &task);
task.iov = iov;
@@ -918,6 +986,7 @@ static int coroutine_fn vitastor_co_pwritev(BlockDriverState *bs,
{
VitastorClient *client = bs->opaque;
VitastorRPC task;
vitastor_co_pin_to_bs_ctx(bs);
vitastor_co_init_task(bs, &task);
task.iov = iov;
@@ -991,6 +1060,7 @@ static int coroutine_fn vitastor_co_block_status(BlockDriverState *bs,
#endif
VitastorRPC task;
VitastorClient *client = bs->opaque;
vitastor_co_pin_to_bs_ctx(bs);
uint64_t inode = client->watch ? vitastor_c_inode_get_num(client->watch) : client->inode;
uint8_t bit = 0;
if (client->last_bitmap && client->last_bitmap_inode == inode &&
@@ -1103,6 +1173,7 @@ static int coroutine_fn vitastor_co_flush(BlockDriverState *bs)
{
VitastorClient *client = bs->opaque;
VitastorRPC task;
vitastor_co_pin_to_bs_ctx(bs);
vitastor_co_init_task(bs, &task);
qemu_mutex_lock(&client->mutex);
@@ -1188,6 +1259,11 @@ static BlockDriver bdrv_vitastor = {
#endif
.bdrv_close = vitastor_close,
// Re-register fd handlers when the bs is moved to a different AioContext
// (live migration, drain, iothread reassignment).
.bdrv_detach_aio_context = vitastor_detach_aio_context,
.bdrv_attach_aio_context = vitastor_attach_aio_context,
// Option list for the create operation
#if QEMU_VERSION_MAJOR >= 3 || QEMU_VERSION_MAJOR == 2 && QEMU_VERSION_MINOR > 0
.create_opts = &vitastor_create_opts,
+52 -6
View File
@@ -62,6 +62,13 @@ const char *help_text =
"All usual Vitastor config options like --config_path <path_to_config> may also be specified in CLI.\n"
;
struct ublk_request
{
uint64_t ublk_cmd;
int index;
int result;
};
class ublk_server
{
protected:
@@ -241,7 +248,7 @@ help:
// Create client
epmgr = new epoll_manager_t(ringloop);
cli = new cluster_client_t(ringloop, epmgr->tfd, cfg);
cli = cluster_client_t::create(ringloop, epmgr->tfd, cfg);
// cli->config contains merged config
if (!cfg["queue_depth"].is_null())
@@ -273,7 +280,7 @@ help:
}
if (!inode)
{
watch = cli->st_cli.watch_inode(image_name);
watch = cli->st_cli->watch_inode(image_name);
device_size = watch->cfg.size;
if (!watch->cfg.num || !device_size)
{
@@ -282,9 +289,9 @@ help:
exit(1);
}
}
const bool writeback = !cli->get_immediate_commit(inode);
auto pool_it = cli->st_cli.pool_config.find(INODE_POOL(inode ? inode : watch->cfg.num));
if (pool_it == cli->st_cli.pool_config.end())
const bool writeback = !cli->get_immediate_commit(inode ? inode : watch->cfg.num);
auto pool_it = cli->st_cli->pool_config.find(INODE_POOL(inode ? inode : watch->cfg.num));
if (pool_it == cli->st_cli->pool_config.end())
{
fprintf(stderr, "Pool %u does not exist\n", INODE_POOL(inode ? inode : watch->cfg.num));
exit(1);
@@ -331,6 +338,11 @@ help:
daemonize_fork(notifyfd);
close(notifyfd[0]);
}
consumer.loop = [this]()
{
submit_postponed();
};
ringloop->register_consumer(&consumer);
start_device(recover);
if (pidfile != "")
write_pid();
@@ -350,6 +362,7 @@ help:
ringloop->wait();
}
cli->flush();
ringloop->unregister_consumer(&consumer);
delete cli;
delete epmgr;
cli = NULL;
@@ -553,6 +566,8 @@ protected:
ublksrv_ctrl_dev_info ublk_dev = {};
ublksrv_io_desc *ublk_queue = NULL;
std::vector<uint8_t*> buffers;
ring_consumer_t consumer;
std::vector<ublk_request> postponed_requests;
void open_control()
{
@@ -734,9 +749,19 @@ protected:
ctrl_fd = -1;
}
void submit_request(uint64_t ublk_cmd, int i, int res)
bool submit_request(uint64_t ublk_cmd, int i, int res)
{
io_uring_sqe *sqe = ringloop->get_sqe();
if (!sqe)
{
// Handle full io_uring by postponing the request
postponed_requests.push_back((ublk_request){
.ublk_cmd = ublk_cmd,
.index = i,
.result = res,
});
return false;
}
ring_data_t* data = ((ring_data_t*)sqe->user_data);
sqe->fd = cdev_fd;
sqe->opcode = IORING_OP_URING_CMD;
@@ -750,6 +775,22 @@ protected:
cmd->addr = (uint64_t)buffers[i];
cmd->result = res;
data->callback = [this, i](ring_data_t *data) { exec_request(data->res, i); };
return true;
}
void submit_postponed()
{
int sent = 0;
while (postponed_requests.size())
{
ublk_request r = postponed_requests.back();
postponed_requests.pop_back();
if (!submit_request(r.ublk_cmd, r.index, r.result))
break;
sent++;
}
if (sent)
ringloop->submit();
}
void exec_request(int res, int i)
@@ -864,6 +905,11 @@ protected:
int sync_ublk_cmd(uint32_t cmd_op, void *addr, uint32_t len, uint16_t dev_path_len = 0, uint64_t data0 = 0)
{
io_uring_sqe *sqe = ringloop->get_sqe();
if (!sqe)
{
fprintf(stderr, "Error: io_uring is full when trying to execute a control command\n");
exit(1);
}
sqe->fd = ctrl_fd;
sqe->opcode = IORING_OP_URING_CMD;
sqe->ioprio = 0;
+1 -1
View File
@@ -6,7 +6,7 @@ includedir=${prefix}/@CMAKE_INSTALL_INCLUDEDIR@
Name: Vitastor
Description: Vitastor client library
Version: 3.0.13
Version: 3.0.15
Libs: -L${libdir} -lvitastor_client
Cflags: -I${includedir}
+13 -13
View File
@@ -103,7 +103,7 @@ vitastor_c *vitastor_c_create_qemu(QEMUSetFDHandler *aio_set_fd_handler, void *a
rdma_device, rdma_port_num, rdma_gid_index, rdma_mtu, log_level
);
auto self = vitastor_c_create_qemu_common(aio_set_fd_handler, aio_context);
self->cli = new cluster_client_t(NULL, self->tfd, cfg_json);
self->cli = cluster_client_t::create(NULL, self->tfd, cfg_json);
return self;
}
@@ -126,7 +126,7 @@ vitastor_c *vitastor_c_create_qemu_uring(QEMUSetFDHandler *aio_set_fd_handler, v
);
auto self = vitastor_c_create_qemu_common(aio_set_fd_handler, aio_context);
self->ringloop = ringloop;
self->cli = new cluster_client_t(self->ringloop, self->tfd, cfg_json);
self->cli = cluster_client_t::create(self->ringloop, self->tfd, cfg_json);
ringloop->loop();
return self;
}
@@ -150,7 +150,7 @@ vitastor_c *vitastor_c_create_uring(const char *config_path, const char *etcd_ho
vitastor_c *self = new vitastor_c;
self->ringloop = ringloop;
self->epmgr = new epoll_manager_t(self->ringloop);
self->cli = new cluster_client_t(self->ringloop, self->epmgr->tfd, cfg_json);
self->cli = cluster_client_t::create(self->ringloop, self->epmgr->tfd, cfg_json);
ringloop->loop();
return self;
}
@@ -191,7 +191,7 @@ vitastor_c *vitastor_c_create_uring_json(const char **options, int options_len)
vitastor_c *self = new vitastor_c;
self->ringloop = ringloop;
self->epmgr = new epoll_manager_t(self->ringloop);
self->cli = new cluster_client_t(self->ringloop, self->epmgr->tfd, cfg_json);
self->cli = cluster_client_t::create(self->ringloop, self->epmgr->tfd, cfg_json);
ringloop->loop();
return self;
}
@@ -206,7 +206,7 @@ vitastor_c *vitastor_c_create_epoll_json(const char **options, int options_len)
json11::Json cfg_json(cfg);
vitastor_c *self = new vitastor_c;
self->epmgr = new epoll_manager_t(NULL);
self->cli = new cluster_client_t(NULL, self->epmgr->tfd, cfg_json);
self->cli = cluster_client_t::create(NULL, self->epmgr->tfd, cfg_json);
return self;
}
@@ -396,7 +396,7 @@ void vitastor_c_watch_inode(vitastor_c *client, char *image, VitastorIOHandler c
{
client->cli->on_ready([=]()
{
auto watch = client->cli->st_cli.watch_inode(std::string(image));
auto watch = client->cli->st_cli->watch_inode(std::string(image));
cb(opaque, (long)watch);
});
if (client->ringloop)
@@ -407,7 +407,7 @@ void vitastor_c_watch_inode(vitastor_c *client, char *image, VitastorIOHandler c
void vitastor_c_close_watch(vitastor_c *client, void *handle)
{
client->cli->st_cli.close_watch((inode_watch_t*)handle);
client->cli->st_cli->close_watch((inode_watch_t*)handle);
}
uint64_t vitastor_c_inode_get_size(void *handle)
@@ -424,8 +424,8 @@ uint64_t vitastor_c_inode_get_num(void *handle)
uint32_t vitastor_c_inode_get_block_size(vitastor_c *client, uint64_t inode_num)
{
auto pool_it = client->cli->st_cli.pool_config.find(INODE_POOL(inode_num));
if (pool_it == client->cli->st_cli.pool_config.end())
auto pool_it = client->cli->st_cli->pool_config.find(INODE_POOL(inode_num));
if (pool_it == client->cli->st_cli->pool_config.end())
return 0;
auto & pool_cfg = pool_it->second;
uint32_t pg_data_size = (pool_cfg.scheme == POOL_SCHEME_REPLICATED ? 1 : pool_cfg.pg_size-pool_cfg.parity_chunks);
@@ -434,8 +434,8 @@ uint32_t vitastor_c_inode_get_block_size(vitastor_c *client, uint64_t inode_num)
uint32_t vitastor_c_inode_get_bitmap_granularity(vitastor_c *client, uint64_t inode_num)
{
auto pool_it = client->cli->st_cli.pool_config.find(INODE_POOL(inode_num));
if (pool_it == client->cli->st_cli.pool_config.end())
auto pool_it = client->cli->st_cli->pool_config.find(INODE_POOL(inode_num));
if (pool_it == client->cli->st_cli->pool_config.end())
return 0;
// FIXME: READ_BITMAP may fails if parent bitmap granularity differs from inode bitmap granularity
return pool_it->second.bitmap_granularity;
@@ -471,8 +471,8 @@ uint64_t vitastor_c_inode_get_mod_revision(void *handle)
uint32_t vitastor_c_inode_get_immediate_commit(vitastor_c *client, uint64_t inode_num)
{
auto pool_it = client->cli->st_cli.pool_config.find(INODE_POOL(inode_num));
if (pool_it == client->cli->st_cli.pool_config.end())
auto pool_it = client->cli->st_cli->pool_config.find(INODE_POOL(inode_num));
if (pool_it == client->cli->st_cli->pool_config.end())
return 0;
return pool_it->second.immediate_commit;
}
+1 -1
View File
@@ -555,7 +555,7 @@ static int run(cli_tool_t *p, json11::Json::object cfg)
json11::Json cfg_j = cfg;
p->ringloop = new ring_loop_t(RINGLOOP_DEFAULT_SIZE);
p->epmgr = new epoll_manager_t(p->ringloop);
p->cli = new cluster_client_t(p->ringloop, p->epmgr->tfd, cfg_j);
p->cli = cluster_client_t::create(p->ringloop, p->epmgr->tfd, cfg_j);
p->loop_and_wait(action_cb, [&](const cli_result_t & r)
{
result = r;
+4 -4
View File
@@ -35,7 +35,7 @@ struct alloc_osd_t
{ "target", "VERSION" },
{ "version", 0 },
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/osd/stats/"+std::to_string(new_id)
parent->cli->st_cli->etcd_prefix+"/osd/stats/"+std::to_string(new_id)
) },
},
} },
@@ -43,7 +43,7 @@ struct alloc_osd_t
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/osd/stats/"+std::to_string(new_id)
parent->cli->st_cli->etcd_prefix+"/osd/stats/"+std::to_string(new_id)
) },
{ "value", base64_encode("{}") },
} },
@@ -52,8 +52,8 @@ struct alloc_osd_t
{ "failure", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/osd/stats/") },
{ "range_end", base64_encode(parent->cli->st_cli.etcd_prefix+"/osd/stats0") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/osd/stats/") },
{ "range_end", base64_encode(parent->cli->st_cli->etcd_prefix+"/osd/stats0") },
{ "keys_only", true },
} },
},
+13 -13
View File
@@ -8,8 +8,8 @@
void cli_tool_t::change_parent(inode_t cur, inode_t new_parent, cli_result_t *result)
{
auto cur_cfg_it = cli->st_cli.inode_config.find(cur);
if (cur_cfg_it == cli->st_cli.inode_config.end())
auto cur_cfg_it = cli->st_cli->inode_config.find(cur);
if (cur_cfg_it == cli->st_cli->inode_config.end())
{
char buf[128];
snprintf(buf, 128, "Inode 0x%jx disappeared", cur);
@@ -18,13 +18,13 @@ void cli_tool_t::change_parent(inode_t cur, inode_t new_parent, cli_result_t *re
}
inode_config_t new_cfg = cur_cfg_it->second;
std::string cur_name = new_cfg.name;
std::string cur_cfg_key = base64_encode(cli->st_cli.etcd_prefix+
std::string cur_cfg_key = base64_encode(cli->st_cli->etcd_prefix+
"/config/inode/"+std::to_string(INODE_POOL(cur))+
"/"+std::to_string(INODE_NO_POOL(cur)));
new_cfg.parent_id = new_parent;
json11::Json::object cur_cfg_json = cli->st_cli.serialize_inode_cfg(&new_cfg);
json11::Json::object cur_cfg_json = cli->st_cli->serialize_inode_cfg(&new_cfg);
waiting++;
cli->st_cli.etcd_txn_slow(json11::Json::object {
cli->st_cli->etcd_txn_slow(json11::Json::object {
{ "compare", json11::Json::array {
json11::Json::object {
{ "target", "MOD" },
@@ -53,8 +53,8 @@ void cli_tool_t::change_parent(inode_t cur, inode_t new_parent, cli_result_t *re
}
else if (new_parent)
{
auto new_parent_it = cli->st_cli.inode_config.find(new_parent);
std::string new_parent_name = new_parent_it != cli->st_cli.inode_config.end()
auto new_parent_it = cli->st_cli->inode_config.find(new_parent);
std::string new_parent_name = new_parent_it != cli->st_cli->inode_config.end()
? new_parent_it->second.name : "<unknown>";
*result = (cli_result_t){
.text = "Parent of layer "+cur_name+" (inode "+std::to_string(INODE_NO_POOL(cur))+
@@ -77,7 +77,7 @@ void cli_tool_t::change_parent(inode_t cur, inode_t new_parent, cli_result_t *re
void cli_tool_t::etcd_txn(json11::Json txn)
{
waiting++;
cli->st_cli.etcd_txn_slow(txn, [this](std::string err, json11::Json res)
cli->st_cli->etcd_txn_slow(txn, [this](std::string err, json11::Json res)
{
waiting--;
if (err != "")
@@ -91,7 +91,7 @@ void cli_tool_t::etcd_txn(json11::Json txn)
inode_config_t* cli_tool_t::get_inode_cfg(const std::string & name)
{
for (auto & ic: cli->st_cli.inode_config)
for (auto & ic: cli->st_cli->inode_config)
{
if (ic.second.name == name)
{
@@ -171,11 +171,11 @@ void cli_tool_t::iterate_kvs_1(json11::Json kvs, const std::string & prefix, std
bool is_pool = prefix == "/pool/stats/";
for (auto & kv_item: kvs.array_items())
{
auto kv = cli->st_cli.parse_etcd_kv(kv_item);
auto kv = cli->st_cli->parse_etcd_kv(kv_item);
uint64_t num = 0;
char null_byte = 0;
// OSD or pool number
int scanned = sscanf(kv.key.substr(cli->st_cli.etcd_prefix.size() + prefix.size()).c_str(), "%ju%c", &num, &null_byte);
int scanned = sscanf(kv.key.substr(cli->st_cli->etcd_prefix.size() + prefix.size()).c_str(), "%ju%c", &num, &null_byte);
if (scanned != 1 || !num || is_pool && num >= POOL_ID_MAX)
{
fprintf(stderr, "Invalid key in etcd: %s\n", kv.key.c_str());
@@ -190,12 +190,12 @@ void cli_tool_t::iterate_kvs_2(json11::Json kvs, const std::string & prefix, std
bool is_inode = prefix == "/config/inode/" || prefix == "/inode/stats/";
for (auto & kv_item: kvs.array_items())
{
auto kv = cli->st_cli.parse_etcd_kv(kv_item);
auto kv = cli->st_cli->parse_etcd_kv(kv_item);
pool_id_t pool_id = 0;
uint64_t num = 0;
char null_byte = 0;
// pool+pg or pool+inode
int scanned = sscanf(kv.key.substr(cli->st_cli.etcd_prefix.size() + prefix.size()).c_str(),
int scanned = sscanf(kv.key.substr(cli->st_cli->etcd_prefix.size() + prefix.size()).c_str(),
"%u/%ju%c", &pool_id, &num, &null_byte);
if (scanned != 2 || !pool_id || is_inode && INODE_POOL(num) || !is_inode && num >= UINT32_MAX)
{
+29 -29
View File
@@ -46,7 +46,7 @@ struct image_creator_t
void loop()
{
auto & pools = parent->cli->st_cli.pool_config;
auto & pools = parent->cli->st_cli->pool_config;
if (state >= 1)
goto resume_1;
if (image_name == "")
@@ -117,7 +117,7 @@ struct image_creator_t
goto resume_2;
else if (state == 3)
goto resume_3;
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (ic.second.name == image_name)
{
@@ -193,7 +193,7 @@ resume_3:
} while (!parent->etcd_result["succeeded"].bool_value());
// Save into inode_config for library users to be able to take it from there immediately
new_cfg.mod_revision = parent->etcd_result["header"]["revision"].uint64_value();
parent->cli->st_cli.insert_inode_config(new_cfg);
parent->cli->st_cli->insert_inode_config(new_cfg);
result = (cli_result_t){
.err = 0,
.text = "Image "+image_name+" created",
@@ -215,8 +215,8 @@ resume_3:
goto resume_3;
else if (state == 4)
goto resume_4;
// FIXME: take all info from etcd requests, not mixed with st_cli.inode_config
for (auto & ic: parent->cli->st_cli.inode_config)
// FIXME: take all info from etcd requests, not mixed with st_cli->inode_config
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (ic.second.name == image_name+"@"+new_snap)
{
@@ -271,7 +271,7 @@ resume_4:
} while (!parent->etcd_result["succeeded"].bool_value());
// Save into inode_config for library users to be able to take it from there immediately
new_cfg.mod_revision = parent->etcd_result["header"]["revision"].uint64_value();
parent->cli->st_cli.insert_inode_config(new_cfg);
parent->cli->st_cli->insert_inode_config(new_cfg);
result = (cli_result_t){
.err = 0,
.text = "Snapshot "+image_name+"@"+new_snap+" created",
@@ -291,7 +291,7 @@ resume_4:
return json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/index/maxid/"+std::to_string(new_pool_id)
parent->cli->st_cli->etcd_prefix+"/index/maxid/"+std::to_string(new_pool_id)
) },
} },
};
@@ -303,13 +303,13 @@ resume_4:
max_id_mod_rev = 0;
if (response["response_range"]["kvs"].array_items().size() > 0)
{
auto kv = parent->cli->st_cli.parse_etcd_kv(response["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(response["response_range"]["kvs"][0]);
new_id = 1+INODE_NO_POOL(kv.value.uint64_value());
max_id_mod_rev = kv.mod_revision;
}
// Also check existing inodes - for the case when some inodes are created without changing /index/maxid
auto ino_it = parent->cli->st_cli.inode_config.lower_bound(INODE_WITH_POOL(new_pool_id+1, 0));
if (ino_it != parent->cli->st_cli.inode_config.begin())
auto ino_it = parent->cli->st_cli->inode_config.lower_bound(INODE_WITH_POOL(new_pool_id+1, 0));
if (ino_it != parent->cli->st_cli->inode_config.begin())
{
ino_it--;
if (INODE_POOL(ino_it->first) == new_pool_id && new_id < 1+INODE_NO_POOL(ino_it->first))
@@ -325,7 +325,7 @@ resume_4:
goto resume_3;
if (!new_pool_id)
{
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (ic.second.name == image_name)
{
@@ -339,7 +339,7 @@ resume_4:
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/index/image/"+image_name
parent->cli->st_cli->etcd_prefix+"/index/image/"+image_name
) },
} },
},
@@ -360,7 +360,7 @@ resume_2:
cfg_mod_rev = idx_mod_rev = 0;
if (parent->etcd_result["responses"][1]["response_range"]["kvs"].array_items().size() == 0)
{
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (ic.second.name == image_name)
{
@@ -377,7 +377,7 @@ resume_2:
{
// FIXME: Parse kvs in etcd_state_client automatically
{
auto kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][1]["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][1]["response_range"]["kvs"][0]);
old_id = INODE_NO_POOL(kv.value["id"].uint64_value());
old_pool_id = (pool_id_t)kv.value["pool_id"].uint64_value();
idx_mod_rev = kv.mod_revision;
@@ -393,7 +393,7 @@ resume_2:
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/config/inode/"+
parent->cli->st_cli->etcd_prefix+"/config/inode/"+
std::to_string(old_pool_id)+"/"+std::to_string(old_id)
) },
} },
@@ -411,7 +411,7 @@ resume_3:
return;
}
{
auto kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
size = kv.value["size"].uint64_value();
new_parent_id = kv.value["parent_id"].uint64_value();
uint64_t parent_pool_id = kv.value["parent_pool"].uint64_value();
@@ -439,7 +439,7 @@ resume_3:
{ "target", "VERSION" },
{ "version", 0 },
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/config/inode/"+
parent->cli->st_cli->etcd_prefix+"/config/inode/"+
std::to_string(new_pool_id)+"/"+std::to_string(new_id)
) },
},
@@ -447,31 +447,31 @@ resume_3:
{ "target", "VERSION" },
{ "version", 0 },
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/index/image/"+image_name+
parent->cli->st_cli->etcd_prefix+"/index/image/"+image_name+
(new_snap != "" ? "@"+new_snap : "")
) },
},
json11::Json::object {
{ "target", "MOD" },
{ "mod_revision", max_id_mod_rev },
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/index/maxid/"+std::to_string(new_pool_id)) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/index/maxid/"+std::to_string(new_pool_id)) },
},
};
json11::Json::array success = json11::Json::array {
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/config/inode/"+
parent->cli->st_cli->etcd_prefix+"/config/inode/"+
std::to_string(new_pool_id)+"/"+std::to_string(new_id)
) },
{ "value", base64_encode(
json11::Json(parent->cli->st_cli.serialize_inode_cfg(&new_cfg)).dump()
json11::Json(parent->cli->st_cli->serialize_inode_cfg(&new_cfg)).dump()
) },
} },
},
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/index/image/"+image_name) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/index/image/"+image_name) },
{ "value", base64_encode(json11::Json(json11::Json::object{
{ "id", new_id },
{ "pool_id", (uint64_t)new_pool_id },
@@ -481,7 +481,7 @@ resume_3:
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/index/maxid/"+
parent->cli->st_cli->etcd_prefix+"/index/maxid/"+
std::to_string(new_pool_id)
) },
{ "value", base64_encode(std::to_string(new_id)) }
@@ -492,7 +492,7 @@ resume_3:
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/index/image/"+
parent->cli->st_cli->etcd_prefix+"/index/image/"+
image_name+(new_snap != "" ? "@"+new_snap : "")
) },
} },
@@ -511,29 +511,29 @@ resume_3:
{ "target", "MOD" },
{ "mod_revision", cfg_mod_rev },
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/config/inode/"+
parent->cli->st_cli->etcd_prefix+"/config/inode/"+
std::to_string(old_pool_id)+"/"+std::to_string(old_id)
) },
});
checks.push_back(json11::Json::object {
{ "target", "MOD" },
{ "mod_revision", idx_mod_rev },
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/index/image/"+image_name) }
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/index/image/"+image_name) }
});
success.push_back(json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/config/inode/"+
parent->cli->st_cli->etcd_prefix+"/config/inode/"+
std::to_string(old_pool_id)+"/"+std::to_string(old_id)
) },
{ "value", base64_encode(
json11::Json(parent->cli->st_cli.serialize_inode_cfg(&snap_cfg)).dump()
json11::Json(parent->cli->st_cli->serialize_inode_cfg(&snap_cfg)).dump()
) },
} },
});
success.push_back(json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/index/image/"+image_name+"@"+new_snap) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/index/image/"+image_name+"@"+new_snap) },
{ "value", base64_encode(json11::Json(json11::Json::object{
{ "id", old_id },
{ "pool_id", (uint64_t)old_pool_id },
+26 -20
View File
@@ -52,19 +52,19 @@ struct dd_in_info_t
in_seekable = true;
if (iimg != "")
{
iwatch = parent->cli->st_cli.watch_inode(iimg);
iwatch = parent->cli->st_cli->watch_inode(iimg);
if (!iwatch->cfg.num)
{
result = (cli_result_t){ .err = ENOENT, .text = "Image "+iimg+" does not exist" };
parent->cli->st_cli.close_watch(iwatch);
parent->cli->st_cli->close_watch(iwatch);
iwatch = NULL;
return;
}
auto pool_it = parent->cli->st_cli.pool_config.find(INODE_POOL(iwatch->cfg.num));
if (pool_it == parent->cli->st_cli.pool_config.end())
auto pool_it = parent->cli->st_cli->pool_config.find(INODE_POOL(iwatch->cfg.num));
if (pool_it == parent->cli->st_cli->pool_config.end())
{
result = (cli_result_t){ .err = ENOENT, .text = "Pool of image "+iimg+" does not exist" };
parent->cli->st_cli.close_watch(iwatch);
parent->cli->st_cli->close_watch(iwatch);
iwatch = NULL;
return;
}
@@ -131,7 +131,7 @@ struct dd_in_info_t
{
if (iimg != "")
{
parent->cli->st_cli.close_watch(iwatch);
parent->cli->st_cli->close_watch(iwatch);
iwatch = NULL;
}
else if (ifile != "")
@@ -163,11 +163,11 @@ struct dd_out_info_t
pool_config_t *find_pool(cli_tool_t *parent, const std::string & name)
{
if (name == "" && parent->cli->st_cli.pool_config.size() == 1)
if (name == "" && parent->cli->st_cli->pool_config.size() == 1)
{
return &parent->cli->st_cli.pool_config.begin()->second;
return &parent->cli->st_cli->pool_config.begin()->second;
}
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
if (pp.second.name == name)
{
@@ -186,14 +186,14 @@ struct dd_out_info_t
if (oimg != "")
{
out_seekable = true;
owatch = parent->cli->st_cli.watch_inode(oimg);
owatch = parent->cli->st_cli->watch_inode(oimg);
if (owatch->cfg.num)
{
auto pool_it = parent->cli->st_cli.pool_config.find(INODE_POOL(owatch->cfg.num));
if (pool_it == parent->cli->st_cli.pool_config.end())
auto pool_it = parent->cli->st_cli->pool_config.find(INODE_POOL(owatch->cfg.num));
if (pool_it == parent->cli->st_cli->pool_config.end())
{
result = (cli_result_t){ .err = ENOENT, .text = "Pool of image "+oimg+" does not exist" };
parent->cli->st_cli.close_watch(owatch);
parent->cli->st_cli->close_watch(owatch);
owatch = NULL;
return true;
}
@@ -209,7 +209,7 @@ struct dd_out_info_t
else
{
result = (cli_result_t){ .err = ENOENT, .text = "Pool to create output image "+oimg+" is not specified" };
parent->cli->st_cli.close_watch(owatch);
parent->cli->st_cli->close_watch(owatch);
owatch = NULL;
return true;
}
@@ -224,14 +224,14 @@ struct dd_out_info_t
if (!out_create)
{
result = (cli_result_t){ .err = ENOENT, .text = "Image "+oimg+" does not exist" };
parent->cli->st_cli.close_watch(owatch);
parent->cli->st_cli->close_watch(owatch);
owatch = NULL;
return true;
}
if (!out_size)
{
result = (cli_result_t){ .err = ENOENT, .text = "Input size is unknown, specify size to create output image "+oimg };
parent->cli->st_cli.close_watch(owatch);
parent->cli->st_cli->close_watch(owatch);
owatch = NULL;
return true;
}
@@ -247,7 +247,7 @@ struct dd_out_info_t
if (!out_size)
{
result = (cli_result_t){ .err = ENOENT, .text = "Input size is unknown, specify size to truncate output image" };
parent->cli->st_cli.close_watch(owatch);
parent->cli->st_cli->close_watch(owatch);
owatch = NULL;
return true;
}
@@ -275,7 +275,7 @@ resume_1:
sub_cb = NULL;
if (result.err)
{
parent->cli->st_cli.close_watch(owatch);
parent->cli->st_cli->close_watch(owatch);
owatch = NULL;
return true;
}
@@ -324,8 +324,14 @@ resume_2:
cluster_op_t *sync_op = new cluster_op_t;
sync_op->opcode = OSD_OP_SYNC;
parent->waiting++;
sync_op->callback = [parent](cluster_op_t *sync_op)
sync_op->callback = [this, parent](cluster_op_t *sync_op)
{
if (sync_op->retval != 0 && !result.err)
{
// Just in case, actually OP_SYNC can't fail
result.err = -sync_op->retval;
result.text = "Failed to sync "+oimg+": "+std::string(strerror(result.err));
}
parent->waiting--;
delete sync_op;
parent->ringloop->wakeup();
@@ -354,7 +360,7 @@ resume_2:
{
if (oimg != "")
{
parent->cli->st_cli.close_watch(owatch);
parent->cli->st_cli->close_watch(owatch);
owatch = NULL;
}
else
+2 -2
View File
@@ -60,7 +60,7 @@ struct cli_describe_t
only_pool = cfg["pool"].uint64_value();
if (!only_pool && cfg["pool"].is_string())
{
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
if (pp.second.name == cfg["pool"].string_value())
{
@@ -125,7 +125,7 @@ struct cli_describe_t
{
uint64_t min_pool = min_inode >> (64-POOL_ID_BITS);
uint64_t max_pool = max_inode >> (64-POOL_ID_BITS);
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
if (pp.first >= min_pool && (!max_pool || pp.first <= max_pool))
{
+2 -2
View File
@@ -134,8 +134,8 @@ struct cli_fix_t
return;
}
auto & obj = objects[processed_count++];
auto pool_cfg_it = parent->cli->st_cli.pool_config.find(INODE_POOL(obj.inode));
if (pool_cfg_it == parent->cli->st_cli.pool_config.end())
auto pool_cfg_it = parent->cli->st_cli->pool_config.find(INODE_POOL(obj.inode));
if (pool_cfg_it == parent->cli->st_cli->pool_config.end())
{
fprintf(stderr, "Object %jx:%jx is from unknown pool\n", obj.inode, obj.stripe);
continue;
+4 -4
View File
@@ -41,8 +41,8 @@ struct snap_flattener_t
chain_list.push_back(cur->num);
while (cur->parent_id != 0 && cur->parent_id != target_cfg->num)
{
auto it = parent->cli->st_cli.inode_config.find(cur->parent_id);
if (it == parent->cli->st_cli.inode_config.end())
auto it = parent->cli->st_cli->inode_config.find(cur->parent_id);
if (it == parent->cli->st_cli->inode_config.end())
{
result = (cli_result_t){
.err = ENOENT,
@@ -97,9 +97,9 @@ struct snap_flattener_t
{ "from", top_parent_name },
{ "to", target_name },
{ "target", target_name },
{ "delete-source", false },
{ "delete_source", false },
{ "cas", use_cas },
{ "fsync-interval", fsync_interval },
{ "fsync_interval", fsync_interval },
});
// Wait for it
resume_1:
+18 -18
View File
@@ -37,7 +37,7 @@ struct image_lister_t
{
if (list_pool_name != "")
{
for (auto & ic: parent->cli->st_cli.pool_config)
for (auto & ic: parent->cli->st_cli->pool_config)
{
if (ic.second.name == list_pool_name)
{
@@ -52,14 +52,14 @@ struct image_lister_t
return;
}
}
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (list_pool_id && INODE_POOL(ic.second.num) != list_pool_id)
{
continue;
}
auto pool_it = parent->cli->st_cli.pool_config.find(INODE_POOL(ic.second.num));
bool good_pool = pool_it != parent->cli->st_cli.pool_config.end();
auto pool_it = parent->cli->st_cli->pool_config.find(INODE_POOL(ic.second.num));
bool good_pool = pool_it != parent->cli->st_cli->pool_config.end();
auto item = json11::Json::object {
{ "name", ic.second.name },
{ "size", ic.second.size },
@@ -73,8 +73,8 @@ struct image_lister_t
};
if (ic.second.parent_id)
{
auto p_it = parent->cli->st_cli.inode_config.find(ic.second.parent_id);
item["parent_name"] = p_it != parent->cli->st_cli.inode_config.end()
auto p_it = parent->cli->st_cli->inode_config.find(ic.second.parent_id);
item["parent_name"] = p_it != parent->cli->st_cli->inode_config.end()
? p_it->second.name : "";
item["parent_pool_id"] = (uint64_t)INODE_POOL(ic.second.parent_id);
item["parent_inode_num"] = INODE_NO_POOL(ic.second.parent_id);
@@ -96,21 +96,21 @@ struct image_lister_t
json11::Json::object {
{ "request_range", (list_pool_id
? json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/pool/stats/"+std::to_string(list_pool_id)) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/pool/stats/"+std::to_string(list_pool_id)) },
}
: json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/pool/stats/") },
{ "range_end", base64_encode(parent->cli->st_cli.etcd_prefix+"/pool/stats0") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/pool/stats/") },
{ "range_end", base64_encode(parent->cli->st_cli->etcd_prefix+"/pool/stats0") },
}) },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/inode/stats"+
parent->cli->st_cli->etcd_prefix+"/inode/stats"+
(list_pool_id ? "/"+std::to_string(list_pool_id) : "")+"/"
) },
{ "range_end", base64_encode(
parent->cli->st_cli.etcd_prefix+"/inode/stats"+
parent->cli->st_cli->etcd_prefix+"/inode/stats"+
(list_pool_id ? "/"+std::to_string(list_pool_id) : "")+"0"
) },
} },
@@ -131,11 +131,11 @@ resume_1:
std::map<pool_id_t, uint64_t> pool_pg_real_size;
for (auto & kv_item: space_info["responses"][0]["response_range"]["kvs"].array_items())
{
auto kv = parent->cli->st_cli.parse_etcd_kv(kv_item);
auto kv = parent->cli->st_cli->parse_etcd_kv(kv_item);
// pool ID
pool_id_t pool_id;
char null_byte = 0;
int scanned = sscanf(kv.key.substr(parent->cli->st_cli.etcd_prefix.length()).c_str(), "/pool/stats/%u%c", &pool_id, &null_byte);
int scanned = sscanf(kv.key.substr(parent->cli->st_cli->etcd_prefix.length()).c_str(), "/pool/stats/%u%c", &pool_id, &null_byte);
if (scanned != 1 || !pool_id || pool_id >= POOL_ID_MAX)
{
fprintf(stderr, "Invalid key in etcd: %s\n", kv.key.c_str());
@@ -146,12 +146,12 @@ resume_1:
}
for (auto & kv_item: space_info["responses"][1]["response_range"]["kvs"].array_items())
{
auto kv = parent->cli->st_cli.parse_etcd_kv(kv_item);
auto kv = parent->cli->st_cli->parse_etcd_kv(kv_item);
// pool ID & inode number
pool_id_t pool_id;
inode_t only_inode_num;
char null_byte = 0;
int scanned = sscanf(kv.key.substr(parent->cli->st_cli.etcd_prefix.length()).c_str(),
int scanned = sscanf(kv.key.substr(parent->cli->st_cli->etcd_prefix.length()).c_str(),
"/inode/stats/%u/%ju%c", &pool_id, &only_inode_num, &null_byte);
if (scanned != 2 || !pool_id || pool_id >= POOL_ID_MAX || INODE_POOL(only_inode_num) != 0)
{
@@ -161,8 +161,8 @@ resume_1:
inode_t inode_num = INODE_WITH_POOL(pool_id, only_inode_num);
uint64_t used_size = kv.value["raw_used"].uint64_value();
// save stats
auto pool_it = parent->cli->st_cli.pool_config.find(pool_id);
if (pool_it != parent->cli->st_cli.pool_config.end())
auto pool_it = parent->cli->st_cli->pool_config.find(pool_id);
if (pool_it != parent->cli->st_cli->pool_config.end())
{
auto & pool_cfg = pool_it->second;
used_size = used_size / (pool_pg_real_size[pool_id] ? pool_pg_real_size[pool_id] : 1)
@@ -176,7 +176,7 @@ resume_1:
{ "size", 0 },
{ "readonly", false },
{ "pool_id", (uint64_t)INODE_POOL(inode_num) },
{ "pool_name", pool_it != parent->cli->st_cli.pool_config.end()
{ "pool_name", pool_it != parent->cli->st_cli->pool_config.end()
? (pool_it->second.name == "" ? "<Unnamed>" : pool_it->second.name) : "?" },
{ "inode_num", INODE_NO_POOL(inode_num) },
{ "inode_id", inode_num },
+13 -7
View File
@@ -110,8 +110,8 @@ struct snap_merger_t
cur->parent_id != to_cfg->num &&
cur->parent_id != 0)
{
auto it = parent->cli->st_cli.inode_config.find(cur->parent_id);
if (it == parent->cli->st_cli.inode_config.end())
auto it = parent->cli->st_cli->inode_config.find(cur->parent_id);
if (it == parent->cli->st_cli->inode_config.end())
{
result = (cli_result_t){
.err = ENOENT,
@@ -166,7 +166,7 @@ struct snap_merger_t
//
// <from> - <layer 1> - <target> - <to>
// \- <layer 2> <---------X-------- NOT ALLOWED
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
auto it = sources.find(ic.second.num);
if (it == sources.end() && ic.second.parent_id != 0)
@@ -182,7 +182,7 @@ struct snap_merger_t
.text = "Layers at or above "+(check_delete_source ? from_name : target_name)+
", but below "+to_name+" are not allowed to have other children, but "+
ic.second.name+" is a child of "+
parent->cli->st_cli.inode_config.at(ic.second.parent_id).name,
parent->cli->st_cli->inode_config.at(ic.second.parent_id).name,
};
state = 100;
return;
@@ -213,7 +213,7 @@ struct snap_merger_t
uint64_t get_block_size(inode_t inode, uint32_t *bitmap_granularity)
{
auto & pool_cfg = parent->cli->st_cli.pool_config.at(INODE_POOL(inode));
auto & pool_cfg = parent->cli->st_cli->pool_config.at(INODE_POOL(inode));
uint64_t pg_data_size = (pool_cfg.scheme == POOL_SCHEME_REPLICATED ? 1 : pool_cfg.pg_size-pool_cfg.parity_chunks);
if (bitmap_granularity)
*bitmap_granularity = pool_cfg.bitmap_granularity;
@@ -251,6 +251,8 @@ struct snap_merger_t
goto resume_100;
// Get parents and so on
start_merge();
if (state == 100)
return;
// First list lower layers
list_errcode.clear();
list_layers(true);
@@ -418,7 +420,7 @@ struct snap_merger_t
}
if (!pgs_left)
{
auto & name = parent->cli->st_cli.inode_config.at(src).name;
auto & name = parent->cli->st_cli->inode_config.at(src).name;
if (list_errcode.find(src) != list_errcode.end())
{
fprintf(stderr, "Failed to get listing of layer %s (inode %ju in pool %u): %s (code %d)\n",
@@ -612,8 +614,10 @@ struct snap_merger_t
subop->offset = offset;
subop->len = 0;
subop->flags = OSD_OP_IGNORE_READONLY | OSD_OP_WAIT_UP_TIMEOUT;
subop->callback = [](cluster_op_t *subop)
in_flight++;
subop->callback = [this](cluster_op_t *subop)
{
in_flight--;
if (subop->retval != 0)
{
fprintf(stderr, "error deleting from layer 0x%jx at offset %jx: %s", subop->inode, subop->offset, strerror(-subop->retval));
@@ -640,8 +644,10 @@ struct snap_merger_t
uint64_t to = last_written_offset;
cluster_op_t *subop = new cluster_op_t;
subop->opcode = OSD_OP_SYNC;
in_flight++;
subop->callback = [this, to](cluster_op_t *subop)
{
in_flight--;
delete subop;
// We can now delete source data between <from> and <to>
// But to do this we have to keep all object lists in memory :-(
+11 -11
View File
@@ -53,7 +53,7 @@ struct image_changer_t
state = 100;
return;
}
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (ic.second.name == image_name)
{
@@ -74,7 +74,7 @@ struct image_changer_t
state = 100;
return;
}
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (ic.second.parent_id == inode_num)
{
@@ -110,9 +110,9 @@ struct image_changer_t
cb = parent->start_rm_data(json11::Json::object {
{ "inode", INODE_NO_POOL(inode_num) },
{ "pool", (uint64_t)INODE_POOL(inode_num) },
{ "fsync-interval", fsync_interval },
{ "min-offset", ((new_size+4095)/4096)*4096 },
{ "down-ok", down_ok },
{ "fsync_interval", fsync_interval },
{ "min_offset", ((new_size+4095)/4096)*4096 },
{ "down_ok", down_ok },
});
resume_1:
while (!cb(result))
@@ -153,7 +153,7 @@ resume_1:
cfg.name = new_name;
}
{
std::string cur_cfg_key = base64_encode(parent->cli->st_cli.etcd_prefix+
std::string cur_cfg_key = base64_encode(parent->cli->st_cli->etcd_prefix+
"/config/inode/"+std::to_string(INODE_POOL(inode_num))+
"/"+std::to_string(INODE_NO_POOL(inode_num)));
checks.push_back(json11::Json::object {
@@ -166,7 +166,7 @@ resume_1:
{ "request_put", json11::Json::object {
{ "key", cur_cfg_key },
{ "value", base64_encode(json11::Json(
parent->cli->st_cli.serialize_inode_cfg(&cfg)
parent->cli->st_cli->serialize_inode_cfg(&cfg)
).dump()) },
} }
});
@@ -174,10 +174,10 @@ resume_1:
if (new_name != "")
{
std::string old_idx_key = base64_encode(
parent->cli->st_cli.etcd_prefix+"/index/image/"+image_name
parent->cli->st_cli->etcd_prefix+"/index/image/"+image_name
);
std::string new_idx_key = base64_encode(
parent->cli->st_cli.etcd_prefix+"/index/image/"+new_name
parent->cli->st_cli->etcd_prefix+"/index/image/"+new_name
);
checks.push_back(json11::Json::object {
{ "target", "MOD" },
@@ -229,9 +229,9 @@ resume_2:
cfg.mod_revision = parent->etcd_result["header"]["revision"].uint64_value();
if (new_name != "")
{
parent->cli->st_cli.inode_by_name.erase(image_name);
parent->cli->st_cli->inode_by_name.erase(image_name);
}
parent->cli->st_cli.insert_inode_config(cfg);
parent->cli->st_cli->insert_inode_config(cfg);
result = (cli_result_t){
.err = 0,
.text = "Image "+image_name+" modified",
+8 -8
View File
@@ -68,12 +68,12 @@ struct osd_changer_t
{ "success", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/osd/stats/"+std::to_string(osd_num)) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/osd/stats/"+std::to_string(osd_num)) },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
} },
},
} },
@@ -89,14 +89,14 @@ resume_1:
return;
}
{
auto osd_stats = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]).value;
auto osd_stats = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]).value;
if (!osd_stats.is_object() && !force)
{
result = (cli_result_t){ .err = ENOENT, .text = "OSD "+std::to_string(osd_num)+" does not exist. Use --force to set configuration anyway" };
state = 100;
return;
}
auto kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][1]["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][1]["response_range"]["kvs"][0]);
osd_cfg_mod_rev = kv.mod_revision;
osd_cfg = kv.value.object_items();
if (set_reweight)
@@ -124,7 +124,7 @@ resume_1:
{
compare.push_back(json11::Json::object {
{ "target", "MOD" },
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
{ "result", "LESS" },
{ "mod_revision", osd_cfg_mod_rev+1 },
});
@@ -133,7 +133,7 @@ resume_1:
{
compare.push_back(json11::Json::object {
{ "target", "VERSION" },
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
{ "version", 0 },
});
}
@@ -141,7 +141,7 @@ resume_1:
{
success.push_back(json11::Json::object {
{ "request_delete_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
} },
});
}
@@ -149,7 +149,7 @@ resume_1:
{
success.push_back(json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/osd/"+std::to_string(osd_num)) },
{ "value", base64_encode(json11::Json(osd_cfg).dump()) },
} },
});
+7 -7
View File
@@ -62,19 +62,19 @@ struct osd_tree_printer_t
{ "success", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/node_placement") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/node_placement") },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/osd/") },
{ "range_end", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/osd0") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/osd/") },
{ "range_end", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/osd0") },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/osd/stats/") },
{ "range_end", base64_encode(parent->cli->st_cli.etcd_prefix+"/osd/stats0") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/osd/stats/") },
{ "range_end", base64_encode(parent->cli->st_cli->etcd_prefix+"/osd/stats0") },
} },
},
} },
@@ -91,7 +91,7 @@ resume_1:
}
for (auto & item: parent->etcd_result["responses"][0]["response_range"]["kvs"].array_items())
{
node_placement = parent->cli->st_cli.parse_etcd_kv(item).value;
node_placement = parent->cli->st_cli->parse_etcd_kv(item).value;
}
parent->iterate_kvs_1(parent->etcd_result["responses"][1]["response_range"]["kvs"], "/config/osd/", [&](uint64_t cur_osd, json11::Json value)
{
@@ -132,7 +132,7 @@ resume_1:
.parent = kv.second["host"].string_value(),
.size = kv.second["size"].uint64_value(),
.free = kv.second["free"].uint64_value(),
.up = parent->cli->st_cli.peer_states.find(kv.first) != parent->cli->st_cli.peer_states.end(),
.up = parent->cli->st_cli->peer_states.find(kv.first) != parent->cli->st_cli->peer_states.end(),
.reweight = 1,
.noout = false,
.block_size = (uint32_t)kv.second["data_block_size"].uint64_value(),
+4 -4
View File
@@ -32,7 +32,7 @@ struct pg_lister_t
if (pool_name != "")
{
pool_id = 0;
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
if (pp.second.name == pool_name)
{
@@ -51,8 +51,8 @@ struct pg_lister_t
{ "success", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/pgstats"+(pool_id ? "/"+std::to_string(pool_id)+"/" : "/")) },
{ "range_end", base64_encode(parent->cli->st_cli.etcd_prefix+"/pgstats"+(pool_id ? "/"+std::to_string(pool_id)+"0" : "0")) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/pgstats"+(pool_id ? "/"+std::to_string(pool_id)+"/" : "/")) },
{ "range_end", base64_encode(parent->cli->st_cli->etcd_prefix+"/pgstats"+(pool_id ? "/"+std::to_string(pool_id)+"0" : "0")) },
} },
},
} },
@@ -129,7 +129,7 @@ resume_1:
}
}
json11::Json::array pgs;
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
if ((!pool_id || pp.first == pool_id) && (pool_name == "" || pp.second.name == pool_name))
{
+21 -21
View File
@@ -60,8 +60,8 @@ struct pool_creator_t
// Validate pool parameters
{
auto new_cfg = cfg.object_items();
result.text = validate_pool_config(new_cfg, json11::Json(), parent->cli->st_cli.global_block_size,
parent->cli->st_cli.global_bitmap_granularity, force);
result.text = validate_pool_config(new_cfg, json11::Json(), parent->cli->st_cli->global_block_size,
parent->cli->st_cli->global_bitmap_granularity, force);
cfg = new_cfg;
}
if (result.text != "")
@@ -72,7 +72,7 @@ struct pool_creator_t
}
// Validate pool name
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
if (pp.second.name == cfg["name"].string_value())
{
@@ -95,13 +95,13 @@ resume_1:
{ "success", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/node_placement") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/node_placement") },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/osd/stats/") },
{ "range_end", base64_encode(parent->cli->st_cli.etcd_prefix+"/osd/stats0") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/osd/stats/") },
{ "range_end", base64_encode(parent->cli->st_cli->etcd_prefix+"/osd/stats0") },
} },
},
} },
@@ -120,7 +120,7 @@ resume_2:
// Get state_node_tree based on node_placement and osd stats
{
auto node_placement_kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
auto node_placement_kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
timespec tv_now;
clock_gettime(CLOCK_REALTIME, &tv_now);
uint64_t osd_out_time = parent->cli->config["osd_out_time"].uint64_value();
@@ -145,7 +145,7 @@ resume_2:
{
osd_configs.push_back(json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/osd/"+osd_num.as_string()) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/osd/"+osd_num.as_string()) },
} }
});
}
@@ -168,7 +168,7 @@ resume_3:
std::vector<json11::Json> osd_configs;
for (auto & ocr: parent->etcd_result["responses"].array_items())
{
auto kv = parent->cli->st_cli.parse_etcd_kv(ocr["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(ocr["response_range"]["kvs"][0]);
osd_configs.push_back(kv.value);
}
state_node_tree = filter_state_node_tree_by_tags(state_node_tree, osd_configs);
@@ -229,7 +229,7 @@ resume_5:
{ "success", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
} }
},
} },
@@ -246,7 +246,7 @@ resume_6:
}
{
// Add new pool
auto kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
new_pools = create_pool(kv);
if (new_pools.is_string())
{
@@ -261,7 +261,7 @@ resume_6:
{ "compare", json11::Json::array {
json11::Json::object {
{ "target", "MOD" },
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
{ "result", "LESS" },
{ "mod_revision", new_pools_mod_rev+1 },
}
@@ -269,7 +269,7 @@ resume_6:
{ "success", json11::Json::array {
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
{ "value", base64_encode(new_pools.dump()) },
} },
},
@@ -307,9 +307,9 @@ resume_8:
parent->waiting++;
parent->epmgr->tfd->set_timer(create_check.interval, false, [this](int timer_id)
{
if (parent->cli->st_cli.pool_config.find(new_id) != parent->cli->st_cli.pool_config.end())
if (parent->cli->st_cli->pool_config.find(new_id) != parent->cli->st_cli->pool_config.end())
{
auto & pool_cfg = parent->cli->st_cli.pool_config[new_id];
auto & pool_cfg = parent->cli->st_cli->pool_config[new_id];
create_check.passed = pool_cfg.real_pg_count > 0;
for (auto pg_it = pool_cfg.pg_config.begin(); pg_it != pool_cfg.pg_config.end(); pg_it++)
{
@@ -459,7 +459,7 @@ resume_8:
osd_bs = UINT32_MAX;
}
}
if (osd_bs && osd_bs != UINT32_MAX && osd_bs != parent->cli->st_cli.global_block_size)
if (osd_bs && osd_bs != UINT32_MAX && osd_bs != parent->cli->st_cli->global_block_size)
{
fprintf(stderr, "Auto-selecting block_size=%s because all pool OSDs use it\n", format_size(osd_bs, false, true).c_str());
upd["block_size"] = osd_bs;
@@ -478,7 +478,7 @@ resume_8:
osd_bg = UINT32_MAX;
}
}
if (osd_bg && osd_bg != UINT32_MAX && osd_bg != parent->cli->st_cli.global_bitmap_granularity)
if (osd_bg && osd_bg != UINT32_MAX && osd_bg != parent->cli->st_cli->global_bitmap_granularity)
{
fprintf(stderr, "Auto-selecting bitmap_granularity=%s because all pool OSDs use it\n", format_size(osd_bg, false, true).c_str());
upd["bitmap_granularity"] = osd_bg;
@@ -498,7 +498,7 @@ resume_8:
osd_imm = UINT32_MAX-1;
}
}
if (osd_imm < UINT32_MAX-1 && osd_imm != parent->cli->st_cli.global_immediate_commit)
if (osd_imm < UINT32_MAX-1 && osd_imm != parent->cli->st_cli->global_immediate_commit)
{
const char *imm_str = osd_imm == IMMEDIATE_NONE ? "none" : (osd_imm == IMMEDIATE_ALL ? "all" : "small");
fprintf(stderr, "Auto-selecting immediate_commit=%s because all pool OSDs use it\n", imm_str);
@@ -530,13 +530,13 @@ resume_8:
block_size = cfg["block_size"].uint64_value()
? cfg["block_size"].uint64_value()
: parent->cli->st_cli.global_block_size;
: parent->cli->st_cli->global_block_size;
bitmap_granularity = cfg["bitmap_granularity"].uint64_value()
? cfg["bitmap_granularity"].uint64_value()
: parent->cli->st_cli.global_bitmap_granularity;
: parent->cli->st_cli->global_bitmap_granularity;
immediate_commit = cfg["immediate_commit"].is_string()
? etcd_state_client_t::parse_immediate_commit(cfg["immediate_commit"].string_value(), IMMEDIATE_ALL)
: parent->cli->st_cli.global_immediate_commit;
: parent->cli->st_cli->global_immediate_commit;
for (auto osd_num_json: state_node_tree["osds"].array_items())
{
+14 -14
View File
@@ -63,37 +63,37 @@ struct pool_lister_t
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/pool/stats/"
parent->cli->st_cli->etcd_prefix+"/pool/stats/"
) },
{ "range_end", base64_encode(
parent->cli->st_cli.etcd_prefix+"/pool/stats0"
parent->cli->st_cli->etcd_prefix+"/pool/stats0"
) },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/osd/stats/"
parent->cli->st_cli->etcd_prefix+"/osd/stats/"
) },
{ "range_end", base64_encode(
parent->cli->st_cli.etcd_prefix+"/osd/stats0"
parent->cli->st_cli->etcd_prefix+"/osd/stats0"
) },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/config/pools"
parent->cli->st_cli->etcd_prefix+"/config/pools"
) },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/config/osd/"
parent->cli->st_cli->etcd_prefix+"/config/osd/"
) },
{ "range_end", base64_encode(
parent->cli->st_cli.etcd_prefix+"/config/osd0"
parent->cli->st_cli->etcd_prefix+"/config/osd0"
) },
} },
},
@@ -113,7 +113,7 @@ resume_1:
auto config_pools = space_info["responses"][2]["response_range"]["kvs"][0];
if (!config_pools.is_null())
{
config_pools = parent->cli->st_cli.parse_etcd_kv(config_pools).value;
config_pools = parent->cli->st_cli->parse_etcd_kv(config_pools).value;
}
parent->iterate_kvs_1(space_info["responses"][0]["response_range"]["kvs"], "/pool/stats/", [&](uint64_t pool_id, json11::Json value)
{
@@ -134,7 +134,7 @@ resume_1:
}
});
// Calculate max_avail for each pool
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
auto & pool_cfg = pp.second;
uint64_t pool_avail = UINT64_MAX;
@@ -238,10 +238,10 @@ resume_1:
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/pgstats/"
parent->cli->st_cli->etcd_prefix+"/pgstats/"
) },
{ "range_end", base64_encode(
parent->cli->st_cli.etcd_prefix+"/pgstats0"
parent->cli->st_cli->etcd_prefix+"/pgstats0"
) },
} },
},
@@ -289,10 +289,10 @@ resume_1:
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/inode/stats/"
parent->cli->st_cli->etcd_prefix+"/inode/stats/"
) },
{ "range_end", base64_encode(
parent->cli->st_cli.etcd_prefix+"/inode/stats0"
parent->cli->st_cli->etcd_prefix+"/inode/stats0"
) },
} },
},
@@ -478,7 +478,7 @@ resume_3:
auto total = st["object_count"].uint64_value();
auto obj_size = st["block_size"].uint64_value();
if (!obj_size)
obj_size = parent->cli->st_cli.global_block_size;
obj_size = parent->cli->st_cli->global_block_size;
if (st["scheme"] == "ec")
obj_size *= st["pg_size"].uint64_value() - st["parity_chunks"].uint64_value();
else if (st["scheme"] == "xor")
+9 -9
View File
@@ -56,7 +56,7 @@ resume_0:
{ "success", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
} }
},
} },
@@ -73,7 +73,7 @@ resume_1:
}
{
// Parse received pools from etcd
auto kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
// Get pool by name or ID
old_cfg = json11::Json();
@@ -103,8 +103,8 @@ resume_1:
// Update pool
new_cfg = cfg;
result.text = validate_pool_config(new_cfg, old_cfg, parent->cli->st_cli.global_block_size,
parent->cli->st_cli.global_bitmap_granularity, force);
result.text = validate_pool_config(new_cfg, old_cfg, parent->cli->st_cli->global_block_size,
parent->cli->st_cli->global_bitmap_granularity, force);
if (result.text != "")
{
result.err = EINVAL;
@@ -115,8 +115,8 @@ resume_1:
if (new_cfg.find("used_for_app") != new_cfg.end() && !force)
{
// Check that pool doesn't have images
auto img_it = parent->cli->st_cli.inode_config.lower_bound(INODE_WITH_POOL(pool_id, 0));
if (img_it != parent->cli->st_cli.inode_config.end() &&
auto img_it = parent->cli->st_cli->inode_config.lower_bound(INODE_WITH_POOL(pool_id, 0));
if (img_it != parent->cli->st_cli->inode_config.end() &&
INODE_POOL(img_it->first) == pool_id &&
new_cfg["used_for_app"].string_value().substr(0, 3) == "fs:" &&
img_it->second.name == new_cfg["used_for_app"].string_value().substr(3))
@@ -124,7 +124,7 @@ resume_1:
// Only allow metadata image to exist in the FS pool
img_it++;
}
if (img_it != parent->cli->st_cli.inode_config.end() && INODE_POOL(img_it->first) == pool_id)
if (img_it != parent->cli->st_cli->inode_config.end() && INODE_POOL(img_it->first) == pool_id)
{
result = (cli_result_t){ .err = ENOENT, .text = "Pool "+pool_name+" has block images, delete them before using it for VitastorFS, S3 or another app" };
state = 100;
@@ -145,7 +145,7 @@ resume_1:
{ "compare", json11::Json::array {
json11::Json::object {
{ "target", "MOD" },
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
{ "result", "LESS" },
{ "mod_revision", pools_mod_rev+1 },
}
@@ -153,7 +153,7 @@ resume_1:
{ "success", json11::Json::array {
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
{ "value", base64_encode(new_pools.dump()) },
} },
},
+7 -7
View File
@@ -56,7 +56,7 @@ struct pool_remover_t
// Get pool id by name (if name given)
if (pool_name != "")
{
for (auto & ic: parent->cli->st_cli.pool_config)
for (auto & ic: parent->cli->st_cli->pool_config)
{
if (ic.second.name == pool_name)
{
@@ -73,7 +73,7 @@ struct pool_remover_t
pool_name = "id " + std::to_string(pool_id);
// Look-up pool id in pool_config
if (parent->cli->st_cli.pool_config.find(pool_id) != parent->cli->st_cli.pool_config.end())
if (parent->cli->st_cli->pool_config.find(pool_id) != parent->cli->st_cli->pool_config.end())
{
pool_valid = 1;
}
@@ -92,7 +92,7 @@ struct pool_remover_t
{
std::string images;
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (pool_id && INODE_POOL(ic.second.num) != pool_id)
{
@@ -124,7 +124,7 @@ resume_1:
{ "success", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
} }
},
} },
@@ -141,7 +141,7 @@ resume_2:
}
{
// Parse received pools from etcd
auto kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
// Remove pool
auto p = kv.value.object_items();
@@ -166,7 +166,7 @@ resume_2:
{ "compare", json11::Json::array {
json11::Json::object {
{ "target", "MOD" },
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
{ "result", "LESS" },
{ "mod_revision", pools_mod_rev+1 },
}
@@ -174,7 +174,7 @@ resume_2:
{ "success", json11::Json::array {
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/config/pools") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/config/pools") },
{ "value", base64_encode(new_pools.dump()) },
} },
},
+3 -3
View File
@@ -82,8 +82,8 @@ struct cli_raw_ls_t
}
if (!pg_count || !pg_stripe_size || !osds.size())
{
auto pool_it = parent->cli->st_cli.pool_config.find(pool_id);
if (pool_it == parent->cli->st_cli.pool_config.end())
auto pool_it = parent->cli->st_cli->pool_config.find(pool_id);
if (pool_it == parent->cli->st_cli->pool_config.end())
{
result = (cli_result_t){ .err = EINVAL, .text = "pg_count, pg_stripe_size and osds are required if the pool does not exist" };
state = 100;
@@ -127,7 +127,7 @@ struct cli_raw_ls_t
for (; osd_pos < osd_list.size() && parent->waiting < parent->parallel_osds; osd_pos++)
{
uint64_t osd_num = osd_list[osd_pos];
if (parent->cli->st_cli.peer_states[osd_num].is_null())
if (parent->cli->st_cli->peer_states[osd_num].is_null())
{
fprintf(stderr, "OSD %ju is unavailable, skipping\n", osd_num);
continue;
+45 -45
View File
@@ -129,7 +129,7 @@ resume_1:
{
if (merge_children[current_child] == inverse_child)
continue;
rebased_images.push_back(parent->cli->st_cli.inode_config.at(merge_children[current_child]).name);
rebased_images.push_back(parent->cli->st_cli->inode_config.at(merge_children[current_child]).name);
start_merge_child(merge_children[current_child], merge_children[current_child]);
resume_2:
while (!wait_result(2))
@@ -176,8 +176,8 @@ resume_6:
if (chain_list[current_child] == inverse_parent)
continue;
{
auto parent_it = parent->cli->st_cli.inode_config.find(chain_list[current_child]);
if (parent_it != parent->cli->st_cli.inode_config.end())
auto parent_it = parent->cli->st_cli->inode_config.find(chain_list[current_child]);
if (parent_it != parent->cli->st_cli->inode_config.end())
deleted_images.push_back(parent_it->second.name);
deleted_ids.push_back(chain_list[current_child]);
}
@@ -264,8 +264,8 @@ resume_100:
chain_list.push_back(cur->num);
while (cur->num != from_cfg->num && cur->parent_id != 0)
{
auto it = parent->cli->st_cli.inode_config.find(cur->parent_id);
if (it == parent->cli->st_cli.inode_config.end())
auto it = parent->cli->st_cli->inode_config.find(cur->parent_id);
if (it == parent->cli->st_cli->inode_config.end())
{
char buf[1024];
snprintf(buf, 1024, "Parent inode of layer %s (id 0x%jx) not found", cur->name.c_str(), cur->parent_id);
@@ -288,7 +288,7 @@ resume_100:
{
sources[item] = i--;
}
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
if (!ic.second.parent_id)
{
@@ -319,7 +319,7 @@ resume_100:
reads.push_back(json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+
parent->cli->st_cli->etcd_prefix+
"/inode/stats/"+std::to_string(INODE_POOL(inode))+
"/"+std::to_string(INODE_NO_POOL(inode))
) },
@@ -332,7 +332,7 @@ resume_100:
reads.push_back(json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+
parent->cli->st_cli->etcd_prefix+
"/inode/stats/"+std::to_string(INODE_POOL(inode))+
"/"+std::to_string(INODE_NO_POOL(inode))
) },
@@ -340,7 +340,7 @@ resume_100:
});
}
parent->waiting++;
parent->cli->st_cli.etcd_txn_slow(json11::Json::object {
parent->cli->st_cli->etcd_txn_slow(json11::Json::object {
{ "success", reads },
}, [this](std::string err, json11::Json data)
{
@@ -357,19 +357,19 @@ resume_100:
{
continue;
}
auto kv = parent->cli->st_cli.parse_etcd_kv(inode_result["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(inode_result["response_range"]["kvs"][0]);
pool_id_t pool_id = 0;
inode_t inode = 0;
char null_byte = 0;
int scanned = sscanf(kv.key.c_str() + parent->cli->st_cli.etcd_prefix.length()+13, "%u/%ju%c", &pool_id, &inode, &null_byte);
int scanned = sscanf(kv.key.c_str() + parent->cli->st_cli->etcd_prefix.length()+13, "%u/%ju%c", &pool_id, &inode, &null_byte);
if (scanned != 2 || !inode)
{
result = (cli_result_t){ .err = EIO, .text = "Bad key returned from etcd: "+kv.key };
state = 100;
return;
}
auto pool_cfg_it = parent->cli->st_cli.pool_config.find(pool_id);
if (pool_cfg_it == parent->cli->st_cli.pool_config.end())
auto pool_cfg_it = parent->cli->st_cli->pool_config.find(pool_id);
if (pool_cfg_it == parent->cli->st_cli->pool_config.end())
{
result = (cli_result_t){ .err = ENOENT, .text = "Pool "+std::to_string(pool_id)+" does not exist" };
state = 100;
@@ -412,8 +412,8 @@ resume_100:
void rename_inverse_parent()
{
auto child_it = parent->cli->st_cli.inode_config.find(inverse_child);
if (child_it == parent->cli->st_cli.inode_config.end())
auto child_it = parent->cli->st_cli->inode_config.find(inverse_child);
if (child_it == parent->cli->st_cli->inode_config.end())
{
char buf[1024];
snprintf(buf, 1024, "Inode 0x%jx disappeared", inverse_child);
@@ -421,8 +421,8 @@ resume_100:
state = 100;
return;
}
auto target_it = parent->cli->st_cli.inode_config.find(inverse_parent);
if (target_it == parent->cli->st_cli.inode_config.end())
auto target_it = parent->cli->st_cli->inode_config.find(inverse_parent);
if (target_it == parent->cli->st_cli->inode_config.end())
{
char buf[1024];
snprintf(buf, 1024, "Inode 0x%jx disappeared", inverse_parent);
@@ -435,17 +435,17 @@ resume_100:
inverse_child_name = child_cfg->name;
inverse_parent_name = target_cfg->name;
std::string child_cfg_key = base64_encode(
parent->cli->st_cli.etcd_prefix+
parent->cli->st_cli->etcd_prefix+
"/config/inode/"+std::to_string(INODE_POOL(inverse_child))+
"/"+std::to_string(INODE_NO_POOL(inverse_child))
);
std::string target_cfg_key = base64_encode(
parent->cli->st_cli.etcd_prefix+
parent->cli->st_cli->etcd_prefix+
"/config/inode/"+std::to_string(INODE_POOL(inverse_parent))+
"/"+std::to_string(INODE_NO_POOL(inverse_parent))
);
std::string target_idx_key = base64_encode(
parent->cli->st_cli.etcd_prefix+"/index/image/"+inverse_parent_name
parent->cli->st_cli->etcd_prefix+"/index/image/"+inverse_parent_name
);
// Fill new configuration
inode_config_t new_cfg = *child_cfg;
@@ -480,12 +480,12 @@ resume_100:
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", target_cfg_key },
{ "value", base64_encode(json11::Json(parent->cli->st_cli.serialize_inode_cfg(&new_cfg)).dump()) },
{ "value", base64_encode(json11::Json(parent->cli->st_cli->serialize_inode_cfg(&new_cfg)).dump()) },
} },
},
json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/index/image/"+child_cfg->name) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/index/image/"+child_cfg->name) },
{ "value", base64_encode(json11::Json({
{ "id", INODE_NO_POOL(inverse_parent) },
{ "pool_id", (uint64_t)INODE_POOL(inverse_parent) },
@@ -494,14 +494,14 @@ resume_100:
},
};
// Reparent children of inverse_child
for (auto & cp: parent->cli->st_cli.inode_config)
for (auto & cp: parent->cli->st_cli->inode_config)
{
if (cp.second.parent_id == child_cfg->num)
{
auto cp_cfg = cp.second;
cp_cfg.parent_id = inverse_parent;
auto cp_key = base64_encode(
parent->cli->st_cli.etcd_prefix+
parent->cli->st_cli->etcd_prefix+
"/config/inode/"+std::to_string(INODE_POOL(cp.second.num))+
"/"+std::to_string(INODE_NO_POOL(cp.second.num))
);
@@ -514,13 +514,13 @@ resume_100:
txn.push_back(json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", cp_key },
{ "value", base64_encode(json11::Json(parent->cli->st_cli.serialize_inode_cfg(&cp_cfg)).dump()) },
{ "value", base64_encode(json11::Json(parent->cli->st_cli->serialize_inode_cfg(&cp_cfg)).dump()) },
} },
});
}
}
parent->waiting++;
parent->cli->st_cli.etcd_txn_slow(json11::Json::object {
parent->cli->st_cli->etcd_txn_slow(json11::Json::object {
{ "compare", cmp },
{ "success", txn },
}, [this](std::string err, json11::Json res)
@@ -550,8 +550,8 @@ resume_100:
void delete_inode_config(inode_t cur)
{
auto cur_cfg_it = parent->cli->st_cli.inode_config.find(cur);
if (cur_cfg_it == parent->cli->st_cli.inode_config.end())
auto cur_cfg_it = parent->cli->st_cli->inode_config.find(cur);
if (cur_cfg_it == parent->cli->st_cli->inode_config.end())
{
char buf[1024];
snprintf(buf, 1024, "Inode 0x%jx disappeared", cur);
@@ -562,12 +562,12 @@ resume_100:
inode_config_t *cur_cfg = &cur_cfg_it->second;
std::string cur_name = cur_cfg->name;
std::string cur_cfg_key = base64_encode(
parent->cli->st_cli.etcd_prefix+
parent->cli->st_cli->etcd_prefix+
"/config/inode/"+std::to_string(INODE_POOL(cur))+
"/"+std::to_string(INODE_NO_POOL(cur))
);
parent->waiting++;
parent->cli->st_cli.etcd_txn_slow(json11::Json::object {
parent->cli->st_cli->etcd_txn_slow(json11::Json::object {
{ "compare", json11::Json::array {
json11::Json::object {
{ "target", "MOD" },
@@ -584,7 +584,7 @@ resume_100:
},
json11::Json::object {
{ "request_delete_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/index/image/"+cur_name) },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/index/image/"+cur_name) },
} },
},
} },
@@ -604,8 +604,8 @@ resume_100:
return;
}
// Modify inode_config for library users to be able to take it from there immediately
parent->cli->st_cli.inode_by_name.erase(cur_name);
parent->cli->st_cli.inode_config.erase(cur);
parent->cli->st_cli->inode_by_name.erase(cur_name);
parent->cli->st_cli->inode_config.erase(cur);
if (parent->progress)
printf("Layer %s deleted\n", cur_name.c_str());
parent->ringloop->wakeup();
@@ -614,8 +614,8 @@ resume_100:
void start_merge_child(inode_t child_inode, inode_t target_inode)
{
auto child_it = parent->cli->st_cli.inode_config.find(child_inode);
if (child_it == parent->cli->st_cli.inode_config.end())
auto child_it = parent->cli->st_cli->inode_config.find(child_inode);
if (child_it == parent->cli->st_cli->inode_config.end())
{
char buf[1024];
snprintf(buf, 1024, "Inode 0x%jx disappeared", child_inode);
@@ -623,8 +623,8 @@ resume_100:
state = 100;
return;
}
auto target_it = parent->cli->st_cli.inode_config.find(target_inode);
if (target_it == parent->cli->st_cli.inode_config.end())
auto target_it = parent->cli->st_cli->inode_config.find(target_inode);
if (target_it == parent->cli->st_cli->inode_config.end())
{
char buf[1024];
snprintf(buf, 1024, "Inode 0x%jx disappeared", target_inode);
@@ -636,16 +636,16 @@ resume_100:
{ "from", from_name },
{ "to", child_it->second.name },
{ "target", target_it->second.name },
{ "delete-source", false },
{ "delete_source", false },
{ "cas", use_cas },
{ "fsync-interval", fsync_interval },
{ "fsync_interval", fsync_interval },
});
}
void start_mark_deleted(inode_t inode)
{
auto ino_it = parent->cli->st_cli.inode_config.find(inode);
if (ino_it == parent->cli->st_cli.inode_config.end())
auto ino_it = parent->cli->st_cli->inode_config.find(inode);
if (ino_it == parent->cli->st_cli->inode_config.end())
{
char buf[1024];
snprintf(buf, 1024, "Inode 0x%jx disappeared", inode);
@@ -665,8 +665,8 @@ resume_100:
void start_delete_source(inode_t inode)
{
auto source = parent->cli->st_cli.inode_config.find(inode);
if (source == parent->cli->st_cli.inode_config.end())
auto source = parent->cli->st_cli->inode_config.find(inode);
if (source == parent->cli->st_cli->inode_config.end())
{
char buf[1024];
snprintf(buf, 1024, "Inode 0x%jx disappeared", inode);
@@ -677,8 +677,8 @@ resume_100:
cb = parent->start_rm_data(json11::Json::object {
{ "inode", inode },
{ "pool", (uint64_t)INODE_POOL(inode) },
{ "fsync-interval", fsync_interval },
{ "down-ok", down_ok },
{ "fsync_interval", fsync_interval },
{ "down_ok", down_ok },
});
}
};
+10 -11
View File
@@ -44,8 +44,8 @@ struct rm_inode_t
void start_delete()
{
auto pool_it = parent->cli->st_cli.pool_config.find(pool_id);
if (pool_it == parent->cli->st_cli.pool_config.end())
auto pool_it = parent->cli->st_cli->pool_config.find(pool_id);
if (pool_it == parent->cli->st_cli->pool_config.end())
{
result = (cli_result_t){ .err = EINVAL, .text = "Pool does not exist" };
state = 100;
@@ -61,12 +61,11 @@ struct rm_inode_t
}
else
{
rm_pg_t *rm = new rm_pg_t((rm_pg_t){
.pg_num = pg_num,
.objects = std::move(objects),
.obj_done = 0,
.synced = !objects.size() || parent->cli->get_immediate_commit(inode),
});
rm_pg_t *rm = new rm_pg_t();
rm->pg_num = pg_num;
rm->objects = std::move(objects);
rm->obj_done = 0;
rm->synced = !rm->objects.size() || parent->cli->get_immediate_commit(inode);
if (min_offset == 0 && max_offset == 0)
{
total_count += rm->objects.size();
@@ -196,8 +195,8 @@ struct rm_inode_t
{
fprintf(stderr, "Warning: some OSDs don't indicate left_on_dead PG OSDs"
" in delete replies, falling back to simpler checks\n");
auto pool_it = parent->cli->st_cli.pool_config.find(pool_id);
if (pool_it != parent->cli->st_cli.pool_config.end())
auto pool_it = parent->cli->st_cli->pool_config.find(pool_id);
if (pool_it != parent->cli->st_cli->pool_config.end())
{
std::set<osd_num_t> all_peers;
for (auto pg_num: fallback_pgs)
@@ -217,7 +216,7 @@ struct rm_inode_t
all_peers.erase(0);
for (auto peer_osd: all_peers)
{
if (parent->cli->st_cli.peer_states[peer_osd].is_null())
if (parent->cli->st_cli->peer_states[peer_osd].is_null())
inactive_osds.insert(peer_osd);
}
}
+12 -12
View File
@@ -68,14 +68,14 @@ struct rm_osd_t
// Check if OSDs are still up
for (auto osd_id: to_remove)
{
if (parent->cli->st_cli.peer_states.find(osd_id) != parent->cli->st_cli.peer_states.end())
if (parent->cli->st_cli->peer_states.find(osd_id) != parent->cli->st_cli->peer_states.end())
{
is_warning = !allow_up;
still_up.push_back(osd_id);
}
}
// Check if OSDs are still used in data distribution
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
// Will OSD deletion make pool incomplete / down / degraded?
bool pool_incomplete = false, pool_down = false, pool_degraded = false;
@@ -189,14 +189,14 @@ struct rm_osd_t
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/pg/config"
parent->cli->st_cli->etcd_prefix+"/pg/config"
) },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/history/last_clean_pgs"
parent->cli->st_cli->etcd_prefix+"/history/last_clean_pgs"
) },
} },
},
@@ -212,10 +212,10 @@ struct rm_osd_t
return;
}
{
auto kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
auto kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][0]["response_range"]["kvs"][0]);
new_pgs = remove_osds_from_pgs(kv);
new_pgs_mod_rev = kv.mod_revision;
kv = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][1]["response_range"]["kvs"][0]);
kv = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][1]["response_range"]["kvs"][0]);
new_clean_pgs = remove_osds_from_pgs(kv);
new_clean_pgs_mod_rev = kv.mod_revision;
}
@@ -235,14 +235,14 @@ struct rm_osd_t
rm_items[i] = json11::Json::object {
{ "request_delete_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+rm_items[i].string_value()
parent->cli->st_cli->etcd_prefix+rm_items[i].string_value()
) },
} },
};
}
if (!new_pgs.is_null())
{
auto pgs_key = base64_encode(parent->cli->st_cli.etcd_prefix+"/pg/config");
auto pgs_key = base64_encode(parent->cli->st_cli->etcd_prefix+"/pg/config");
rm_items.push_back(json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", pgs_key },
@@ -258,7 +258,7 @@ struct rm_osd_t
}
if (!new_clean_pgs.is_null())
{
auto pgs_key = base64_encode(parent->cli->st_cli.etcd_prefix+"/history/last_clean_pgs");
auto pgs_key = base64_encode(parent->cli->st_cli->etcd_prefix+"/history/last_clean_pgs");
rm_items.push_back(json11::Json::object {
{ "request_put", json11::Json::object {
{ "key", pgs_key },
@@ -375,7 +375,7 @@ struct rm_osd_t
goto resume_0;
history_updates.clear();
history_checks.clear();
for (auto & pp: parent->cli->st_cli.pool_config)
for (auto & pp: parent->cli->st_cli->pool_config)
{
bool update_pg_history = false;
auto & pool_cfg = pp.second;
@@ -420,7 +420,7 @@ struct rm_osd_t
if (update_pg_history)
{
std::string history_key = base64_encode(
parent->cli->st_cli.etcd_prefix+"/pg/history/"+
parent->cli->st_cli->etcd_prefix+"/pg/history/"+
std::to_string(pool_cfg.id)+"/"+std::to_string(pg_num)
);
auto hist = json11::Json::object {
@@ -440,7 +440,7 @@ struct rm_osd_t
{ "target", "MOD" },
{ "key", history_key },
{ "result", "LESS" },
{ "mod_revision", parent->cli->st_cli.etcd_watch_revision_pg+1 },
{ "mod_revision", parent->cli->st_cli->etcd_watch_revision_pg+1 },
});
}
}
+9 -9
View File
@@ -48,7 +48,7 @@ struct wildcard_remover_t
{
auto child_id = ino_it->first;
auto parent_id = ino_it->second;
auto & parent_cfg = parent->cli->st_cli.inode_config.at(parent_id);
auto & parent_cfg = parent->cli->st_cli->inode_config.at(parent_id);
if (parent_cfg.parent_id)
{
auto chain_it = chains.find(parent_cfg.parent_id);
@@ -71,7 +71,7 @@ struct wildcard_remover_t
std::vector<inode_rev_t> ver_chain;
do
{
auto & inode_cfg = parent->cli->st_cli.inode_config.at(child_id);
auto & inode_cfg = parent->cli->st_cli->inode_config.at(child_id);
ver_chain.push_back((inode_rev_t){ .inode_num = child_id, .meta_rev = inode_cfg.mod_revision });
if (child_id == parent_id)
break;
@@ -89,7 +89,7 @@ struct wildcard_remover_t
do
{
rank++;
cur_id = parent->cli->st_cli.inode_config.at(cur_id).parent_id;
cur_id = parent->cli->st_cli->inode_config.at(cur_id).parent_id;
} while (cur_id && cur_id != parent_id);
ranks[parent_id] = rank;
}
@@ -111,7 +111,7 @@ struct wildcard_remover_t
state = 0;
chains.clear();
// Select images to delete
for (auto & ic: parent->cli->st_cli.inode_config)
for (auto & ic: parent->cli->st_cli->inode_config)
{
for (auto & glob: globs)
{
@@ -131,11 +131,11 @@ struct wildcard_remover_t
// Check for parallel changes
for (auto & irev: versioned_chains[i])
{
auto inode_it = parent->cli->st_cli.inode_config.find(irev.inode_num);
if (inode_it == parent->cli->st_cli.inode_config.end() ||
auto inode_it = parent->cli->st_cli->inode_config.find(irev.inode_num);
if (inode_it == parent->cli->st_cli->inode_config.end() ||
inode_it->second.mod_revision > irev.meta_rev)
{
if (inode_it != parent->cli->st_cli.inode_config.end())
if (inode_it != parent->cli->st_cli->inode_config.end())
fprintf(stderr, "Warning: image %s modified by someone else during deletion, restarting wildcard deletion\n", inode_it->second.name.c_str());
else
fprintf(stderr, "Warning: inode %jx modified by someone else during deletion, retrying wildcard deletion\n", irev.inode_num);
@@ -144,8 +144,8 @@ struct wildcard_remover_t
}
// Delete
{
auto from_cfg = parent->cli->st_cli.inode_config.at(versioned_chains[i].back().inode_num);
auto to_cfg = parent->cli->st_cli.inode_config.at(versioned_chains[i].front().inode_num);
auto from_cfg = parent->cli->st_cli->inode_config.at(versioned_chains[i].back().inode_num);
auto to_cfg = parent->cli->st_cli->inode_config.at(versioned_chains[i].front().inode_num);
sub_cfg = cfg.object_items();
sub_cfg.erase("globs");
sub_cfg.erase("exact");
+16 -16
View File
@@ -37,14 +37,14 @@ struct status_printer_t
goto resume_2;
// etcd states
{
auto addrs = parent->cli->st_cli.get_addresses();
auto addrs = parent->cli->st_cli->get_addresses();
etcd_states.resize(addrs.size());
for (int i = 0; i < etcd_states.size(); i++)
{
parent->waiting++;
parent->cli->st_cli.etcd_call_oneshot(
parent->cli->st_cli->etcd_call_oneshot(
addrs[i], "/maintenance/status", json11::Json::object(),
parent->cli->st_cli.etcd_quick_timeout, [this, i](std::string err, json11::Json res)
parent->cli->st_cli->etcd_quick_timeout, [this, i](std::string err, json11::Json res)
{
parent->waiting--;
etcd_states[i] = err != "" ? json11::Json::object{ { "error", err } } : res;
@@ -62,23 +62,23 @@ resume_1:
{ "success", json11::Json::array {
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/mon/") },
{ "range_end", base64_encode(parent->cli->st_cli.etcd_prefix+"/mon0") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/mon/") },
{ "range_end", base64_encode(parent->cli->st_cli->etcd_prefix+"/mon0") },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(
parent->cli->st_cli.etcd_prefix+"/osd/stats/"
parent->cli->st_cli->etcd_prefix+"/osd/stats/"
) },
{ "range_end", base64_encode(
parent->cli->st_cli.etcd_prefix+"/osd/stats0"
parent->cli->st_cli->etcd_prefix+"/osd/stats0"
) },
} },
},
json11::Json::object {
{ "request_range", json11::Json::object {
{ "key", base64_encode(parent->cli->st_cli.etcd_prefix+"/stats") },
{ "key", base64_encode(parent->cli->st_cli->etcd_prefix+"/stats") },
} },
},
} },
@@ -97,7 +97,7 @@ resume_2:
auto osd_stats = parent->etcd_result["responses"][1]["response_range"]["kvs"];
if (parent->etcd_result["responses"][2]["response_range"]["kvs"].array_items().size() > 0)
{
agg_stats = parent->cli->st_cli.parse_etcd_kv(parent->etcd_result["responses"][2]["response_range"]["kvs"][0]).value;
agg_stats = parent->cli->st_cli->parse_etcd_kv(parent->etcd_result["responses"][2]["response_range"]["kvs"][0]).value;
}
int etcd_alive = 0;
uint64_t etcd_db_size = 0;
@@ -120,8 +120,8 @@ resume_2:
std::string mon_master;
for (int i = 0; i < mon_members.size(); i++)
{
auto kv = parent->cli->st_cli.parse_etcd_kv(mon_members[i]);
kv.key = kv.key.substr(parent->cli->st_cli.etcd_prefix.size());
auto kv = parent->cli->st_cli->parse_etcd_kv(mon_members[i]);
kv.key = kv.key.substr(parent->cli->st_cli->etcd_prefix.size());
if (kv.key.substr(0, 12) == "/mon/member/")
mon_count++;
else if (kv.key == "/mon/master")
@@ -153,8 +153,8 @@ resume_2:
osds_nearfull++;
}
}
auto peer_it = parent->cli->st_cli.peer_states.find(stat_osd_num);
if (peer_it != parent->cli->st_cli.peer_states.end())
auto peer_it = parent->cli->st_cli->peer_states.find(stat_osd_num);
if (peer_it != parent->cli->st_cli->peer_states.end())
{
osd_up++;
if (value["slow_ops_primary"].uint64_value() > 0)
@@ -177,7 +177,7 @@ resume_2:
std::string backfillfull_pool_names;
std::map<std::string, int> pgs_by_state;
std::string pgs_by_state_str;
for (auto & pool_pair: parent->cli->st_cli.pool_config)
for (auto & pool_pair: parent->cli->st_cli->pool_config)
{
auto & pool_cfg = pool_pair.second;
bool active = pool_cfg.real_pg_count > 0;
@@ -262,7 +262,7 @@ resume_2:
std::string str(obj_states[i]);
uint64_t obj_n = agg_stats["object_bytes"][str].uint64_value();
if (!obj_n)
obj_n = agg_stats["object_counts"][str].uint64_value() * parent->cli->st_cli.global_block_size;
obj_n = agg_stats["object_counts"][str].uint64_value() * parent->cli->st_cli->global_block_size;
json_status[str+"_data"] = obj_n;
}
printf("%s\n", json11::Json(json_status).dump().c_str());
@@ -275,7 +275,7 @@ resume_2:
std::string str(obj_states[i]);
uint64_t obj_n = agg_stats["object_bytes"][str].uint64_value();
if (!obj_n)
obj_n = agg_stats["object_counts"][str].uint64_value() * parent->cli->st_cli.global_block_size;
obj_n = agg_stats["object_counts"][str].uint64_value() * parent->cli->st_cli->global_block_size;
if (!i || obj_n > 0)
more_states += format_size(obj_n)+" "+str+", ";
}
+20
View File
@@ -317,9 +317,13 @@ int main(int argc, char *argv[])
else
{
// First argument is an OSD device - take metadata layout parameters from it
if (self.options["io"] != "")
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
if (self.dump_load_check_superblock(self.dsk.journal_device))
return 1;
}
if (self.options["io"] != "")
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
return self.dump_journal();
}
else if (!strcmp(cmd[0], "write-journal"))
@@ -340,6 +344,8 @@ int main(int argc, char *argv[])
else
{
// First argument is an OSD device - take metadata layout parameters from it
if (self.options["io"] != "")
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
if (self.dump_load_check_superblock(self.new_journal_device))
return 1;
self.new_journal_device = self.dsk.journal_device;
@@ -366,6 +372,10 @@ int main(int argc, char *argv[])
{
self.dsk.csum_block_size = stoull_full(self.options["csum_block_size"], 0);
}
if (self.options["io"] != "")
{
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
}
return self.write_json_journal(entries);
}
else if (!strcmp(cmd[0], "dump-meta"))
@@ -382,6 +392,8 @@ int main(int argc, char *argv[])
{
// First argument is an OSD device - take metadata layout parameters from it
self.dsk.meta_device = cmd[1];
if (self.options["io"] != "")
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
if (self.dump_load_check_superblock(self.dsk.meta_device))
return 1;
}
@@ -398,6 +410,8 @@ int main(int argc, char *argv[])
self.dsk.calc_lengths();
self.dsk.close_all();
}
if (self.options["io"] != "")
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
return self.dump_meta();
}
else if (!strcmp(cmd[0], "write-meta"))
@@ -412,6 +426,8 @@ int main(int argc, char *argv[])
{
// First argument is an OSD device - take metadata layout parameters from it
self.new_meta_device = cmd[1];
if (self.options["io"] != "")
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
if (self.dump_load_check_superblock(self.new_meta_device))
return 1;
self.new_meta_device = self.dsk.meta_device;
@@ -422,6 +438,8 @@ int main(int argc, char *argv[])
{
// Parse all OSD options from cmdline
self.dsk.parse_config(self.options);
if (self.options["io"] != "")
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
self.dsk.open_data();
self.dsk.open_meta();
self.dsk.open_journal();
@@ -438,6 +456,8 @@ int main(int argc, char *argv[])
fprintf(stderr, "Invalid JSON: %s\n", json_err.c_str());
return 1;
}
if (self.options["io"] != "")
self.dsk.data_io = self.dsk.meta_io = self.dsk.journal_io = self.options["io"];
return self.write_json_meta(meta);
}
else if (!strcmp(cmd[0], "resize"))
+13 -15
View File
@@ -477,11 +477,15 @@ int disk_tool_t::write_json_journal(json11::Json entries)
uint16_t type = t_it->second;
if (type == JE_START)
continue;
uint32_t offset = (uint32_t)rec["offset"].uint64_value();
uint32_t len = (uint32_t)rec["len"].uint64_value();
uint32_t data_csum_blocks = !dsk.data_csum_type || !len ? 0 :
(((offset + len - 1)/dsk.csum_block_size - offset/dsk.csum_block_size + 1));
uint32_t data_csum_size = data_csum_blocks*(dsk.data_csum_type & 0xFF);
uint32_t entry_size = (type == JE_START
? sizeof(journal_entry_start)
: (type == JE_SMALL_WRITE || type == JE_SMALL_WRITE_INSTANT
? sizeof(journal_entry_small_write) + dsk.clean_entry_bitmap_size +
(dsk.data_csum_type ? rec["len"].uint64_value()/dsk.csum_block_size*(dsk.data_csum_type & 0xFF) : 0)
? sizeof(journal_entry_small_write) + dsk.clean_entry_bitmap_size + data_csum_size
: (type == JE_BIG_WRITE || type == JE_BIG_WRITE_INSTANT
? sizeof(journal_entry_big_write) + dsk.clean_entry_bitmap_size +
(dsk.data_csum_type ? rec["len"].uint64_value()/dsk.csum_block_size*(dsk.data_csum_type & 0xFF) : 0)
@@ -507,7 +511,7 @@ int disk_tool_t::write_json_journal(json11::Json entries)
journal_entry *ne = (journal_entry*)(new_journal_ptr + new_journal_in_pos);
if (type == JE_SMALL_WRITE || type == JE_SMALL_WRITE_INSTANT)
{
if (new_journal_data - new_journal_buf + ne->small_write.len > new_journal_len)
if (new_journal_data - new_journal_buf + len > new_journal_len)
{
fprintf(stderr, "Error: entries don't fit to the new journal\n");
free(new_journal_buf);
@@ -523,15 +527,12 @@ int disk_tool_t::write_json_journal(json11::Json entries)
.stripe = sscanf_json(NULL, rec["stripe"]),
},
.version = rec["ver"].uint64_value(),
.offset = (uint32_t)rec["offset"].uint64_value(),
.len = (uint32_t)rec["len"].uint64_value(),
.offset = offset,
.len = len,
.data_offset = (uint64_t)(new_journal_data-new_journal_buf),
.crc32_data = !dsk.data_csum_type ? 0 : (uint32_t)sscanf_json("%x", rec["data_crc32"]),
};
uint32_t data_csum_blocks = !dsk.data_csum_type ? 0 :
(((ne->small_write.offset+ne->small_write.len)/dsk.csum_block_size - ne->small_write.len/dsk.csum_block_size));
uint32_t data_csum_size = data_csum_blocks*(dsk.data_csum_type & 0xFF);
fromhexstr(rec["bitmap"].string_value(), dsk.clean_entry_bitmap_size, ((uint8_t*)ne) + sizeof(journal_entry_small_write) + data_csum_size);
fromhexstr(rec["bitmap"].string_value(), dsk.clean_entry_bitmap_size, ((uint8_t*)ne) + sizeof(journal_entry_small_write));
fromhexstr(rec["data"].string_value(), ne->small_write.len, new_journal_data);
if (ne->small_write.len > 0 && !rec["data"].is_string())
{
@@ -545,7 +546,7 @@ int disk_tool_t::write_json_journal(json11::Json entries)
ne->small_write.crc32_data = crc32c(0, new_journal_data, ne->small_write.len);
else if (dsk.data_csum_type == BLOCKSTORE_CSUM_CRC32C)
{
uint32_t *block_csums = (uint32_t*)(((uint8_t*)ne) + sizeof(journal_entry_small_write));
uint32_t *block_csums = (uint32_t*)(((uint8_t*)ne) + sizeof(journal_entry_small_write) + dsk.clean_entry_bitmap_size);
for (uint32_t i = 0; i < data_csum_blocks; i++)
{
uint32_t block_begin = (ne->small_write.offset/dsk.csum_block_size + i) * dsk.csum_block_size;
@@ -574,12 +575,9 @@ int disk_tool_t::write_json_journal(json11::Json entries)
.len = (uint32_t)rec["len"].uint64_value(),
.location = sscanf_json(NULL, rec["loc"]),
};
uint32_t data_csum_blocks = !dsk.data_csum_type ? 0 :
(((ne->small_write.offset+ne->small_write.len)/dsk.csum_block_size - ne->small_write.len/dsk.csum_block_size));
uint32_t data_csum_size = data_csum_blocks*(dsk.data_csum_type & 0xFF);
fromhexstr(rec["bitmap"].string_value(), dsk.clean_entry_bitmap_size, ((uint8_t*)ne) + sizeof(journal_entry_big_write) + data_csum_size);
fromhexstr(rec["bitmap"].string_value(), dsk.clean_entry_bitmap_size, ((uint8_t*)ne) + sizeof(journal_entry_big_write));
if (dsk.data_csum_type)
fromhexstr(rec["block_csums"].string_value(), data_csum_size, ((uint8_t*)ne) + sizeof(journal_entry_big_write));
fromhexstr(rec["block_csums"].string_value(), data_csum_size, ((uint8_t*)ne) + sizeof(journal_entry_big_write) + dsk.clean_entry_bitmap_size);
}
else if (type == JE_STABLE || type == JE_ROLLBACK || type == JE_DELETE)
{
+16 -3
View File
@@ -95,6 +95,12 @@ close_error:
journal_pos += read_len;
}
}
dsk.meta_format = hdr->version;
dsk.data_block_size = hdr->data_block_size;
dsk.csum_block_size = hdr->csum_block_size;
dsk.data_csum_type = hdr->data_csum_type;
dsk.bitmap_granularity = hdr->bitmap_granularity;
dsk.clean_entry_bitmap_size = (hdr->data_block_size / hdr->bitmap_granularity + 7) / 8;
blockstore_heap_t *heap = new blockstore_heap_t(&dsk, buffer_area, log_level);
// Load heap and just iterate it in memory
hdr_fn(hdr);
@@ -575,7 +581,7 @@ int disk_tool_t::write_json_meta(json11::Json meta)
{
if (new_data_csum_size)
{
fromhexstr(e["data_csum"].string_value(), new_data_csum_size,
fromhexstr(e["block_csums"].string_value(), new_data_csum_size,
((uint8_t*)new_entry) + sizeof(clean_disk_entry) + 2*new_clean_entry_bitmap_size);
}
uint32_t *new_entry_csum = (uint32_t*)(((uint8_t*)new_entry) + new_clean_entry_size - 4);
@@ -616,6 +622,12 @@ int disk_tool_t::write_json_heap(json11::Json meta, json11::Json journal)
new_data_csum_size = (new_meta_hdr->csum_block_size
? ((new_meta_hdr->data_block_size+new_meta_hdr->csum_block_size-1)/new_meta_hdr->csum_block_size*(new_meta_hdr->data_csum_type & 0xFF))
: 0);
dsk.meta_format = new_meta_hdr->version;
dsk.data_block_size = new_meta_hdr->data_block_size;
dsk.csum_block_size = new_meta_hdr->csum_block_size;
dsk.data_csum_type = new_meta_hdr->data_csum_type;
dsk.bitmap_granularity = new_meta_hdr->bitmap_granularity;
dsk.clean_entry_bitmap_size = (new_meta_hdr->data_block_size / new_meta_hdr->bitmap_granularity + 7) / 8;
new_journal_buf = NULL;
if (new_journal_len)
{
@@ -696,7 +708,7 @@ close_err0:
wr->entry_type = wr_type | (write_entry["stable"].bool_value() ? BS_HEAP_STABLE : 0);
wr->lsn = write_entry["lsn"].uint64_value();
wr->version = write_entry["version"].uint64_value();
wr->size = wr->get_size(&heap);
wr->size = wr_size;
if (wr_type == BS_HEAP_SMALL_WRITE || wr_type == BS_HEAP_INTENT_WRITE)
{
wr->small().offset = wr_offset;
@@ -736,6 +748,7 @@ close_err0:
bi.offset = wr_offset;
bi.len = wr_len;
}
wr->size = wr->get_size(&heap);
if (write_entry["bitmap"].is_string() && wr->get_int_bitmap(&heap))
{
fromhexstr(write_entry["bitmap"].string_value(), new_clean_entry_bitmap_size, wr->get_int_bitmap(&heap));
@@ -794,7 +807,7 @@ close_err:
fromhexstr(meta_entry["bitmap"].string_value(), new_clean_entry_bitmap_size, wr->get_int_bitmap(&heap));
fromhexstr(meta_entry["ext_bitmap"].string_value(), new_clean_entry_bitmap_size, wr->get_ext_bitmap(&heap));
if (new_data_csum_size)
fromhexstr(meta_entry["data_csum"].string_value(), new_data_csum_size, wr->get_checksums(&heap));
fromhexstr(meta_entry["block_csums"].string_value(), new_data_csum_size, wr->get_checksums(&heap));
wr->crc32c = wr->calc_crc32c();
assert((uint8_t*)wr + wr->size == new_meta_buf + meta_offset + used_space);
auto j_it = journal_by_object.find(oid);
+3 -6
View File
@@ -213,7 +213,8 @@ void disk_tool_t::resize_init(blockstore_meta_header_v3_t *hdr)
data_idx_diff = ((int64_t)(dsk.data_offset-new_data_offset))/((int64_t)dsk.data_block_size);
free_first = new_data_offset > dsk.data_offset ? (new_data_offset-dsk.data_offset) / dsk.data_block_size : 0;
free_last = (new_data_offset+new_data_len < dsk.data_offset+dsk.data_len)
? (dsk.data_offset+dsk.data_len-new_data_offset-new_data_len) / dsk.data_block_size
? ((dsk.data_offset-new_data_offset)/dsk.data_block_size +
dsk.data_len/dsk.data_block_size - new_data_len/dsk.data_block_size)
: 0;
uint32_t new_clean_entry_header_size = sizeof(clean_disk_entry) + 4 /*entry_csum*/;
new_clean_entry_bitmap_size = dsk.data_block_size / (hdr ? hdr->bitmap_granularity : 4096) / 8;
@@ -575,15 +576,11 @@ int disk_tool_t::resize_rebuild_meta()
new_meta_hdr->bitmap_granularity = dsk.bitmap_granularity ? dsk.bitmap_granularity : 4096;
new_meta_hdr->data_csum_type = dsk.data_csum_type;
new_meta_hdr->csum_block_size = dsk.csum_block_size;
new_meta_hdr->completed_lsn = hdr->completed_lsn;
new_meta_hdr->completed_lsn = hdr ? hdr->completed_lsn : 0;
new_meta_hdr->meta_area_size = new_meta_len;
new_meta_hdr->header_csum = 0;
new_meta_hdr->header_csum = crc32c(0, new_meta_hdr, new_meta_hdr->version == BLOCKSTORE_META_FORMAT_HEAP
? sizeof(blockstore_meta_header_v3_t) : sizeof(blockstore_meta_header_v2_t));
if (hdr->version == BLOCKSTORE_META_FORMAT_HEAP && new_meta_format != BLOCKSTORE_META_FORMAT_HEAP)
{
build_journal_start();
}
},
[&](blockstore_heap_t *heap, heap_entry_t *obj, uint32_t meta_block_num)
{
+2 -2
View File
@@ -146,7 +146,7 @@ void kv_cli_t::run()
// Create client
ringloop = new ring_loop_t(512);
epmgr = new epoll_manager_t(ringloop);
cli = new cluster_client_t(ringloop, epmgr->tfd, cfg);
cli = cluster_client_t::create(ringloop, epmgr->tfd, cfg);
db = new vitastorkv_dbw_t(cli);
// Load image metadata
while (!cli->is_ready())
@@ -527,7 +527,7 @@ void kv_cli_t::handle_cmd(const std::vector<std::string> & cmd, std::function<vo
{
inode_id = 0;
name = trim(name);
for (auto & ic: cli->st_cli.inode_config)
for (auto & ic: cli->st_cli->inode_config)
{
if (ic.second.name == name)
{
+3 -3
View File
@@ -233,7 +233,7 @@ static std::string read_string(uint8_t *data, int size, int *pos)
}
uint32_t len = *(uint32_t*)(data+*pos);
*pos += sizeof(uint32_t);
if (*pos+len > size)
if (len > size-*pos)
{
*pos = -1;
return "";
@@ -510,8 +510,8 @@ void kv_db_t::open(inode_t inode_id, json11::Json cfg, std::function<void(int)>
cb(-EINVAL);
return;
}
auto pool_it = cli->st_cli.pool_config.find(INODE_POOL(inode_id));
if (pool_it == cli->st_cli.pool_config.end())
auto pool_it = cli->st_cli->pool_config.find(INODE_POOL(inode_id));
if (pool_it == cli->st_cli->pool_config.end())
{
cb(-EINVAL);
return;
+1 -1
View File
@@ -272,7 +272,7 @@ void kv_test_t::run(json11::Json cfg)
// Create client
ringloop = new ring_loop_t(512);
epmgr = new epoll_manager_t(ringloop);
cli = new cluster_client_t(ringloop, epmgr->tfd, cfg);
cli = cluster_client_t::create(ringloop, epmgr->tfd, cfg);
db = new vitastorkv_dbw_t(cli);
// Load image metadata
while (!cli->is_ready())

Some files were not shown because too many files have changed in this diff Show More