Picture a Friday night, somewhere past midnight, when an on-call engineer’s phone starts buzzing.
A small SaaS team shipped a new release earlier that day. An alert says latency on one endpoint is climbing, and nobody can tell whether a database has slowed down, a dependency has fallen over, or a disk is nearly full. By the time someone opens the log page and scrolls far enough, the cause turns out to be a config change that multiplied the retry traffic.
The postmortem, though, is about more than the incident. It surfaces something the team should have dealt with earlier. After they dropped their cloud monitoring contract, all they had left were scattered container logs and a single Grafana dashboard.
It wasn’t that they refused to pay. The bill had simply grown frightening. With per-host and per-gigabyte pricing stacked together, every new machine pushed the monitoring cost up with it. When finance asked why monitoring cost more than the servers, they started looking at running it themselves.
That’s where the trouble actually began. Self-hosting observability is easy to install and unpleasant to get wrong. OpenObserve and SigNoz are the two open source projects people mention most, and plenty of comparisons exist online, but most stop at feature lists. For a team on a tight budget, a feature list is readable. The pitfalls are not.
So this piece won’t re-list whose feature set is bigger. The question I care about is where these two stacks will take your daily work if you have one decent machine, one part-time maintainer, and a budget you can’t promise will hold steady next month.
You aren’t self-hosting a tool, you’re signing a long-term maintenance contract
The first mistake in a lot of selection processes is treating open source software as free software.
Neither OpenObserve nor SigNoz charges a license fee. That part is true. What self-hosting really consumes is operations time. Someone has to watch the machine, someone has to clear the disk when it fills, someone has to track version upgrades, and when a bug shows up, someone has to dig through the issue tracker.
Start with the foundations, because the two projects differ more than you’d expect.
OpenObserve is written in Rust, and its pitch is “one binary and you’re running.” Its architecture documentation is blunt about the tradeoffs. The default single-node mode is SQLite plus local disk, meant for light use and testing. If you want durability and elasticity, you can switch to single-node with object storage, writing data as parquet files into S3, MinIO, or Azure Blob. Above that, the high-availability mode needs Kubernetes, object storage, PostgreSQL for metadata, NATS for cluster coordination, and at least one node each of five types: Router, Ingester, Compactor, Querier, and Scheduler (see the OpenObserve architecture docs). Single-node is the main path, and the cluster is a separate road for heavy traffic.
SigNoz takes the more standard cloud-native route. It is essentially OpenTelemetry plus ClickHouse: applications send data through the OTel Collector, the data lands in the columnar ClickHouse database, and a packaged SigNoz binary handles the frontend, the query API, alert rules, and alertmanager (see SigNoz architecture). The upside is alignment with the wider tooling standard. OTLP is an industry-wide protocol, so the way you instrument today survives a future tool change. The cost is accepting ClickHouse as part of your stack. It has real demands on memory and disk IO, and scaling or tuning it feels more like administering a database than configuring a monitoring tool.
My read on the difference: OpenObserve hides its complexity inside “you can start small.” SigNoz puts its complexity on display as “you’ll need to understand this eventually.”
Neither is better in the abstract. The complexity small teams fear most is the first kind: you only wanted to read logs, and instead you’re learning to run a database. The reverse also holds. If you already run on Kubernetes and someone on the team knows ClickHouse, the SigNoz path will feel smoother.
Logs, metrics, and traces: don’t wire all three up on day one
The second common mistake is trying to cover all three pillars of observability immediately.
The complete set is logs, metrics, and traces, and their integration costs are wildly different.
Logs are the easiest to start with. Your application already writes them. Point the output somewhere new, or attach a collection agent, and data arrives. Logs are also the most direct source of information during an incident. The problem facing Lao Zhou’s team, back in that opening scene, would have been visible in the logs.
Metrics come next. CPU, memory, request volume, error rate: a stock Prometheus collector can gather all of it with almost no code changes. Metrics earn their keep by showing you a trend before something breaks, rather than documenting the damage afterward.
Traces are the expensive one. You have to instrument your code, understand how spans and traces relate, and the data volume is large enough that storing it costs real money. Traces answer one specific question: a request crossed eight services, and you need to know which segment was slow. That’s valuable when you run enough services to make pinpointing the slow segment difficult.
My suggestion is three steps, spaced a week or two apart. Start with logs alone: get collection, querying, and retention working, and confirm the team actually opens them. Then add metrics, wiring in basic host and application health, and configure two or three alerts that someone will really respond to. Only then consider traces, starting from one critical business path rather than instrumenting everything.
That order isn’t arbitrary. It maps to your first instinct. If your reflex in an incident is to open the logs, make logs good first. If you’ve started wanting a warning before things go wrong, add metrics. If you’re already losing sleep over slow queries and cross-service calls, add traces. The all-at-once approach usually leaves you with three half-installed pillars, none of them pleasant to use, plus a storage bill you didn’t plan for.
One detail gets overlooked here: retention. Many teams default to keeping everything, the log pile grows, and disk alerts start firing more often than business alerts. A saner approach is to tier by value. Keep the last week of logs available for instant search, and push older data to cold storage or keep only aggregated results. The same logic applies to metrics. Lowering the sample rate isn’t automatically cheaper. Sample too aggressively and every curve flattens out exactly when you need it, which leaves you monitoring nothing.
What actually consumes the budget is storage and human time
Self-hosting has two cost centers. One is machines and storage. The other is you.
OpenObserve has a practical design decision here. It writes data as parquet files that can sit directly in object storage. Object storage costs far less per unit than block storage, which means you can retain logs for a long time cheaply. The official docs also mention that, in the vendor’s own Apple M2 test, the default single-node configuration wrote about 31MB per second, roughly 2.6TB a day (see the OpenObserve architecture docs). Those numbers came from their environment, so don’t plug them straight into your capacity plan, but they do show that single-node isn’t a toy. The tradeoff deserves saying too: buying durability with object storage adds a little read latency compared with local disk, so query-heavy workloads need their own judgment call.
On the SigNoz side, compression is where ClickHouse shines, and columnar storage handles structured data like logs and traces well. The tradeoff is ClickHouse’s sensitivity to memory. The official install docs give a modest floor: running single-node Docker Compose requires at least 4GB of memory allocated to Docker (see SigNoz self-hosted install). That number is useful because it means an 8GB machine can get you started. But once log volume climbs, memory becomes the first bottleneck, and you find yourself reading up on sharding, replicas, and disk selection, which were exactly the topics you hoped to avoid.
Human time is harder to price. From what I’ve seen, day-to-day maintenance of OpenObserve in single-node mode mostly means keeping the disk from filling and not falling behind on upgrades. SigNoz adds a layer of “understand how ClickHouse is behaving.” If someone on the team already enjoys databases, that isn’t a burden. If nobody does, it’s work that has to be scheduled.
Don’t overlook backups and upgrade windows either. Self-hosting means you decide when to upgrade, and it also means you have to stop the service at some point to do it. One reality that never gets counted properly: if the dashboard won’t load after an upgrade, there’s no vendor support line to call, only your own rollback. Exporting configuration on a schedule and keeping a copy of your data are part of self-hosting, not an optional extra.
There’s also a version detail worth knowing. Since v0.130.0, SigNoz no longer maintains the old install.sh and the Docker Compose file in its repository, and the project has moved to Foundry, a declarative installation tool (see SigNoz self-hosted install). Changes like this are routine in open source. They don’t change whether you can self-host, but they do change which tutorial you follow. An older article you find through search may open by telling you to run a script that’s already out of date.
Before you install anything, estimate your data volume
Another mistake Lao Zhou’s team made was discovering, only after the install, that they had no idea how many logs they produced daily.
This belongs before tool selection, not after. The method is simple. Look at how many log lines an average service writes per day, multiply by the average size of a line, and add up every service. If your logs embed full request and response bodies, a single line can easily reach several kilobytes, and the total can run tens of times higher than logs that only record key fields.
That estimate decides almost every downstream choice. A team producing a few hundred MB a day can run on one small machine with local disk and has no reason to touch a cluster. At a few hundred GB a day, single-node will eventually buckle, and the options become object storage or a path like SigNoz with ClickHouse behind it. Storage and retention are the same calculation. Whether you plan to keep seven days or ninety can differ by a factor of ten in cost.
One shortcut: sample the noisiest services first, or drop their log level and turn off debug. Plenty of teams run short on disk not because business logs are heavy but because a few services keep printing debug lines nobody reads.
Self-hosted open source doesn’t mean every feature is free
One more thing slips past people: the open source edition and the enterprise edition cover different ground.
OpenObserve and SigNoz both have open source cores, and both offer commercial or cloud-hosted editions. The community editions usually cover the main capabilities of logs, metrics, and traces, while features that lean toward enterprise operations, such as single sign-on, fine-grained permissions, and auditing, tend to live in the commercial editions. Open source projects need revenue, so this arrangement is normal.
What you want to avoid is choosing from the marketing page and discovering after installation that the feature you need sits in a different edition. My advice is unglamorous but effective. Write a short list of the features your team will actually use, then check the docs to confirm each one exists in the open source edition. For the one or two that matter most, verify them on a test machine before committing.
There’s a side benefit to this exercise: it shows you how much you actually need. Many teams finish the list and realize that their daily work comes down to log search, a few alerts, and a dashboard. The rest of the features are there in case of “someday,” and they aren’t a reason to choose one tool over another today.
So which one do you pick
If you want a single sentence: for most small teams with tight budgets and limited hands, starting with OpenObserve in single-node mode will be less trouble. If you already run Kubernetes and OpenTelemetry deeply, and someone on the team is willing to maintain ClickHouse, SigNoz raises your long-term ceiling.
A little more detail.
Choose OpenObserve when you want one machine, one command, and one dashboard that shows logs and metrics, and you want observability running now. You also want storage costs under control and are willing to put data in object storage, and you aren’t planning a complex cluster in the near term. Its single-node mode lets you upgrade later without buying every part on day one.
Choose SigNoz when your services already report over OTLP, or you plan to do distributed tracing properly. You have a Kubernetes environment and don’t want to maintain a separate, mismatched stack just for monitoring. You accept spending a few extra days learning ClickHouse’s temperament in exchange for a stack closer to the industry standard.
Neither choice is final. In observability, tools get replaced, and what you want to keep stable is the data format and the collection method. That’s why it’s worth standardizing on OpenTelemetry even if you pick OpenObserve today. When you switch backends, you change the receiving end, not your application code.
A thirty-day plan for Lao Zhou’s team
Back to that five-person team from the opening.
They didn’t roll out all three pillars at once. In week one, they connected only application logs to a single-node OpenObserve running on an 8GB server, with data going to object storage and retention set to thirty days. In week two, people started checking the error log each day and annotating a few common errors along the way. In week three, they added a Prometheus collector and wrote three alerts on CPU, memory, and endpoint error rate, alerts that actually notify someone. They left traces alone for now, because most of the product still lives in one monolith and cross-service pinpointing isn’t needed yet.
A month in, Lao Zhou said it didn’t feel like swapping tools. It felt like finally having somewhere to ask questions. Incidents still happen, but at least they know where to look for clues instead of guessing at midnight.
Not every team should self-host. If your services are small and nobody is willing to touch operations, paying a bit more for a hosted service and spending that time on the product is often the better trade. Self-hosting is a tradeoff, and it only makes sense when you have the time to spend on it.
This is what self-hosted observability really looks like. It won’t stop incidents or cut the bill to zero overnight, but it does give you a sense of control: the data sits on your own machines, the cost is predictable, and when something breaks, you have something in hand.
If you’re staring at a monitoring bill right now and hesitating, make one small decision first. Pick a stack, turn on logs only, and run it for two weeks. If that works, talk about metrics and traces. The worst outcome comes from trying to do everything at once, whichever tool you pick. And whether to self-host or buy has never answered itself inside the tool. It depends on how much time you’re willing to set aside for it.



